SenseTime Introduces NEO-unify for End-to-End Multimodal Understanding and Generation

Abstract neural network integrating pixel and text data streams into a unified core, symbolizing multimodal AI.

SenseTime has unveiled NEO-unify, a new end-to-end native architecture designed to integrate multimodal understanding and generation. The company detailed this development in a recent technical blog post titled "NEO-unify: Building Native Multimodal Unified Models End to End."

Traditional multimodal models typically employ a "visual encoder (VE) for understanding, variational autoencoder (VAE) for generation" framework. While effective, this approach separates perception from creation, leading to challenges in module coordination and efficiency. NEO-unify aims to address this by allowing AI to directly process pixels and text within a single architecture, collaboratively performing both understanding and generation tasks.

Diagram comparing traditional separate AI components (VE, VAE) with a single, integrated NEO-unify model.

Diagram comparing traditional separate AI components (VE, VAE) with a single, integrated NEO-unify model.

The new model abandons the use of traditional VEs and VAEs, establishing an end-to-end unified model that processes raw pixels and text. This design has shown improvements in training and computational efficiency while maintaining semantic understanding and detail recovery capabilities.

Addressing Multimodal Architecture Limitations

Multimodal research has long relied on a paradigm where a Vision Encoder (VE) handles perception and understanding, and a Variational Autoencoder (VAE) manages content generation. Although some recent efforts have explored shared encoders, these often introduce new structural trade-offs.

SenseTime, in collaboration with Nanyang Technological University, developed NEO-unify as a native, unified, end-to-end multimodal model architecture. This new paradigm bypasses existing debates on visual representation and aims to overcome limitations associated with pre-training priors and scaling law bottlenecks. A key feature is its operation without requiring either a VE or a VAE. SenseTime indicated that more models and open-source results are forthcoming.

Abstract image of a digital barrier dissolving to reveal a clear, interconnected data network, symbolizing architectural breakthrough.

Abstract image of a digital barrier dissolving to reveal a clear, interconnected data network, symbolizing architectural breakthrough.

NEO-unify's Integrated Architecture

NEO-unify represents a step toward a truly end-to-end unified framework that learns directly from nearly lossless information input. The architecture incorporates a nearly lossless visual interface to unify image input and output representations. It also utilizes a native Mixture-of-Transformer (MoT) architecture, which enables understanding and generation to collaborate within the same system.

The model achieves cross-modal training through a unified learning framework. Text is optimized using an autoregressive cross-entropy objective, while vision is optimized through pixel flow matching.

Digital interface showing a Mixture-of-Transformer (MoT) architecture with text and vision optimization flows.

Digital interface showing a Mixture-of-Transformer (MoT) architecture with text and vision optimization flows.

Technical Discoveries and Performance

NEO-unify's encoder-free design allows it to retain both abstract semantics and fine-grained representations simultaneously. Building on previous work, NEO (Diao et al., ICLR 2026), which demonstrated that native end-to-end models can learn rich semantic representations, researchers observed that the independent generation branch could extract and restore fine-grained visual details even with the understanding branch frozen.

Based on this, a 2-billion-parameter (2B) version of NEO-unify was trained. After 90,000 pre-training steps, the model achieved a PSNR of 31.56 and an SSIM of 0.85 on MS COCO 2017. For comparison, Flux VAE recorded metrics of 32.65 and 0.91. These results suggest that native input can support both high-quality semantic understanding and pixel-level detail fidelity without relying on pre-trained VEs or VAEs.

In image editing tasks, NEO-unify (2B) demonstrated robust capabilities even with the understanding branch frozen, significantly reducing the number of input image tokens. After 60,000 steps of mixed training using open-source generation and image editing datasets, the model scored 3.32 on the ImgEdit benchmark.

The encoder-free architecture and MoT backbone in NEO-unify are designed to minimize intrinsic conflicts. When both understanding and generation branches are pre-trained, NEO-unify is jointly trained using the same mid-term training (MT) and supervised fine-tuning (SFT) data. Understanding capabilities remained stable, and generation capabilities converged rapidly, even at lower data ratios and loss weights.

Furthermore, NEO-unify exhibits higher data training efficiency compared to models like Bagel. After web-scale pre-training, followed by MT and SFT on diverse and high-quality data, NEO-unify achieved better performance with fewer training tokens.

Abstract visualization of two data vortices, one compact and intense (NEO-unify), the other larger and diffuse (other models), symbolizing training efficiency.

Abstract visualization of two data vortices, one compact and intense (NEO-unify), the other larger and diffuse (other models), symbolizing training efficiency.

Future Outlook

SenseTime views this architectural exploration as a step toward next-generation intelligent forms, including closed-loop perception and generation, full-modal reasoning, visual reasoning, spatial intelligence, and world models. The company anticipates a roadmap where models can natively think across modalities, moving beyond connecting disparate systems to building unified intelligent agents.

ToolMesh
ToolMesh Weekly

Stay Ahead of the AI Curve

Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.

No spam, unsubscribe at any time.