World Labs introduced Atlas as the first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Pretrained from scratch as a multimodal autoregressive diffusion transformer, it turns a few photos into explorable spaces, 1440p camera-controlled video, and photorealistic RGB-D for robot simulation. Early access opens in the coming weeks.
Key Takeaways
- βOne architecture covers generation, explicit 3D reconstruction, and space-time simulation
- βA handful of ordinary cameras can become a bullet-time studio or a robot RGB-D simulator
- βAtlas beats top open-source reconstruction models; early access is opening soon
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.