On September 1, World Labs, co-founded by Stanford professor Fei-Fei Li, unveiled its next-generation world model, Atlas. The company calls it the world’s first multimodal world model capable of generating images and videos with pixel-level precision, while also enabling 3D scene reconstruction and spatiotemporal simulation.

Fei-Fei Li posted the original text
An All-in-One Architecture for Multimodal Data
Atlas is a multimodal autoregressive diffusion Transformer pretrained from scratch. It can natively process text, images, videos, camera poses, and 3D depth maps.
Its key innovation is to anchor all inputs in 3D space, creating a unified “spatial context.” The model then generates subsequent content based on that context, helping maintain consistency across three-dimensional environments.
Four Core Capabilities Across Generation, Reconstruction, and Simulation
Atlas supports camera-controlled video generation. With just one to six reference images and a predefined camera path, it can generate videos at up to 1440p resolution and one minute in length, while allowing users to precisely control camera angles and movements.
For 3D reconstruction, Atlas can faithfully recreate real-world scenes from just two or three ordinary photographs, outputting explicit geometric data such as point clouds and 3D Gaussian splats. On public benchmarks including DTU, it reportedly outperforms specialized 3D reconstruction models.
System Simulation Test (Image Source: WorldLabs)
Atlas also supports spatiotemporal simulation. It can reconstruct 3D environments from ordinary smartphone videos and generate the RGB images and depth data required for robotic simulation, providing a potentially low-cost and scalable new approach to “Real-to-Sim” training.
In addition, Atlas can generate images and 360-degree panoramas from text prompts.
Targeting Robotics Simulation and Creative Industries
Fei-Fei Li has previously described world models in three categories: renderers, simulators, and planners, emphasizing that simulators are fundamental to helping AI understand the physical world.
Atlas represents a practical implementation of that vision, positioning itself as a spatial intelligence foundation model that can be integrated into applications ranging from robotics and visual effects to game development.
World Labs says Atlas will power future versions of its 3D world-generation product, Marble, and early access has already been opened to selected partners.
The launch comes as spatial intelligence moves from research toward commercialization. World Labs’ recent revenue growth and its acquisition of robotics simulation company SceniX further signal that the company is accelerating its strategic push into the space.