Recommendable! Very impressive!
"World Labs was founded by AI pioneer Fei-Fei Li alongside Justin Johnson, Ben Mildenhall, and Christoph Lassner" (About)
"Fei-Fei Li’s World Labs released Atlas, a new model trained to handle text, images, video, and 3D data in one architecture.
It generates images and video with camera control, producing up to one minute of 1440p video, and reconstructs real scenes into 3D outputs like point clouds and Gaussian splats from as few as two or three input photos. Atlas is a multimodal autoregressive diffusion transformer (a model that generates step by step while also using diffusion techniques common in image generators) and borrows serving techniques from large language models and video diffusion models. ..."
"... Today [9/1/2026] we are introducing Atlas, our next-generation world model. Atlas is an omni model that we pretrained from scratch to natively operate on text, images, video, and 3D.
It is a multimodal autoregressive diffusion transformer: all inputs are combined into a shared spatial context. Atlas uses that context to generate what comes next, staying consistent in 3D with everything it has seen and imagining what lies beyond it. Atlas is built to scale: its performance improves with increased training compute, and we expect this trend to hold as we continue scaling.
Atlas can perform a broad range of tasks spanning world generation, reconstruction, and simulation:
- Camera-Controlled Generation: Atlas generates images and videos from one or more images with pixel-perfect camera control, outputting up to 1 minute of video at 1440p.
- Spatial Reconstruction: Atlas reconstructs real world scenes from one to dozens of input images. It generates both image frames from novel views and explicit 3D outputs, outperforming state-of-the-art models specialized for 3D reconstruction.
- Space-Time Simulation: Atlas models space and time from input videos, reframing videos for dramatic visual effects and enabling Real-to-Sim workflows for robotics.
- Image Generation: Atlas generates images and 360 panoramas from text; it can follow complex prompts, render text, and generate a wide variety of visual styles.
..."
Li Fei Fei
No comments:
Post a Comment