The NEO-unify architecture adopted by
SenseTime Open Sources SenseNova U1, Achieving a Multimodal Native Unified Architecture


The NEO-unify architecture adopted by
SenseTime's AI office agent "Little Raccoon" mobile app is live on iOS and Android. Users long-press voice to command, release to send, and can view/share results remotely, enabling phone commands with PC execution. Feedback survey earns 1000 points; agent features multimodal capabilities.....
SenseNova U1.5-Lite-Preview, an open-source unified multimodal lightweight model by SenseTime, systematically iterates on U1 architecture, integrating visual understanding, reasoning, generation, and editing. With only 8B-MoE parameters, it supports 4K resolution, finer textures, and complex visual control, advancing multimodal capabilities in a compact size.....
SenseTime releases the open-source lightweight multimodal model SenseNova U1.5-Lite-Preview, based on the original NEO-Unify architecture. With only 8B-MoT parameters, it achieves generation and editing quality comparable to closed-source models. It supports multimodal understanding, reasoning, and 4K high resolution, challenging the large model arms race with its compact size.
SenseTime open-sources SenseNova-Vision-7B-MoT, a multi-task vision model integrating object detection, OCR, depth estimation, normal estimation, image segmentation, and multi-view processing in a 7B architecture, providing an efficient base for visual understanding and GUI agents.....
The competition in large models is shifting towards agents. SenseTime is developing the industry's first natively multimodal agent base, integrating a unified core of "understanding, generation, and action", directly benchmarking against GPT-Image 2, and pushing AI from passive Q&A to active execution.....