The NEO-unify architecture adopted by
SenseTime Open Sources SenseNova U1, Achieving a Multimodal Native Unified Architecture


The NEO-unify architecture adopted by
SenseNova U1.5-Lite-Preview, an open-source unified multimodal lightweight model by SenseTime, systematically iterates on U1 architecture, integrating visual understanding, reasoning, generation, and editing. With only 8B-MoE parameters, it supports 4K resolution, finer textures, and complex visual control, advancing multimodal capabilities in a compact size.....
SenseTime releases the open-source lightweight multimodal model SenseNova U1.5-Lite-Preview, based on the original NEO-Unify architecture. With only 8B-MoT parameters, it achieves generation and editing quality comparable to closed-source models. It supports multimodal understanding, reasoning, and 4K high resolution, challenging the large model arms race with its compact size.
SenseTime open-sources SenseNova-Vision-7B-MoT, a multi-task vision model integrating object detection, OCR, depth estimation, normal estimation, image segmentation, and multi-view processing in a 7B architecture, providing an efficient base for visual understanding and GUI agents.....
The competition in large models is shifting towards agents. SenseTime is developing the industry's first natively multimodal agent base, integrating a unified core of "understanding, generation, and action", directly benchmarking against GPT-Image 2, and pushing AI from passive Q&A to active execution.....
SenseTime is secretly developing the multimodal large model U1Pro, targeting design scenarios, led by Chief Scientist Lin Dahua. The model belongs to the "Ri Ri Xin" family, aiming to compete with OpenAI's GPT-Image2, emphasizing long-range logic and thinking capabilities, and expected to launch internal testing and commercial use in July.