The "AI mother" Li Feifei, who is highly regarded in the field of artificial intelligence, founded World Labs, which has recently officially launched the world's first multimodal world model - Atlas. The biggest breakthrough of this model is its ability to generate images and video frames with pixel-level precise camera control and perfectly reconstruct them in 3D space.

As a new model that was pre-trained from scratch, Atlas can seamlessly accept multimodal inputs including camera movements and convert them into stereoscopic 3D views. By precisely locating multiple input views within the model's spatial context, users can generate images and video frames with absolute control, easily achieving the "bullet time" shooting effects like those in Hollywood movies. In addition, it can output clear 3D models from a single or multiple input images, surpassing many top open-source reconstruction models in terms of reconstruction quality.

image.png

In terms of specific working mechanisms, Atlas generates one or more reference images that perfectly match the original content and geometry based on any specified camera position and angle set by the user, smoothly expanding the scene and autonomously imagining and filling in the details of the scenes that are not visible in the input.

According to the official introduction, Atlas is capable of handling a wide range of tasks such as world generation, reconstruction, and simulation. First, in terms of camera-controlled generation, it can output high-definition videos up to 1 minute long at 1440p resolution from one or more images through pixel-precise camera control. Second, in terms of spatial reconstruction, it not only can reconstruct real scenes from dozens of input images, but also generate image frames from novel viewpoints and provide explicit 3D outputs. In addition, the model also supports spatiotemporal simulation functions, modeling space and time through input videos, re-composing videos to enhance dramatic visual effects, and supporting real-to-simulated workflows for robots. At the same time, it fully supports generating images and 360-degree panoramas from text, accurately following complex prompts, rendering text, and presenting various visual styles.

Currently, World Labs has stated that some partners have already obtained early access to the model and plan to gradually open it to a broader group of early adopters in the coming weeks.