Contents

Artificial Intelligence
EU AI Act consulting: system classification, policies, AI governance, training.
Discover →DataGovern
Governance of compliance documentation: policies, evidence and registers kept together and up to date.
Discover DataGovern →On 28 September AMD announced a definitive agreement to acquire World Labs, the lab founded in early 2024 by Fei-Fei Li, whose Stanford group created ImageNet, together with Justin Johnson, Christoph Lassner and Ben Mildenhall. The deal is worth about $8.2 billion, entirely in stock, and is expected to close by the end of 2026, subject to approvals. Li joins AMD as executive vice president and chief scientist, reporting directly to Lisa Su, and the team will keep working on model research.
World Labs builds world models: models that generate, reconstruct and simulate interactive 3D environments from text, images and video, which the company also uses to train and evaluate robots. The problem this class of models tries to address is almost forty years old.
The paradox language models have not solved
In 1988, in Mind Children, robotics researcher Hans Moravec wrote that “it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility”.
The models that write code and solve mathematical problems today have moved the first half of that sentence a long way and the second half very little. Much of the reason lies in data. A language model learns from a corpus of text that already existed, produced by decades of human writing. For physical interaction there is no equivalent corpus: every useful example costs machine time on a real robot, objects to reset, failures to recover from. According to World Labs, not even the video available online covers configurations, materials, robot types and failure conditions systematically.
The second bottleneck is evaluation. To know whether a control policy works, it has to be tried on hardware, and for World Labs this is what makes robot development orders of magnitude slower than language model development. The tasks the company uses in its own demonstrations give a measure of the problem: inserting the end of an elastic cable into a hole, pulling pencils one at a time out of a crowded container, packing a box with two arms. These are movements a person makes without thinking, and they remain hard for a robot to generalise.
What a world model is
The term is used loosely. In June World Labs proposed a functional taxonomy that separates three different things:
- a renderer outputs observations, meaning pixels meant for human eyes, and is judged on visual fidelity: video models and interactive systems such as Google DeepMind’s Genie 3 fall here;
- a simulator outputs state, meaning a representation of the world that is faithful in geometry, physics and dynamics, which both people and programs can compute on;
- a planner outputs actions: given an observation and a goal, it decides what to do next, as vision-language-action models do.
World Labs argues that the simulator is the least discussed of the three and the most important, because it is where an agent can act, learn and be evaluated. The idea of an internal model that anticipates the consequences of actions has a long history: the taxonomy itself traces it back to the “small-scale models” of reality proposed by Kenneth Craik in 1943. Anticipatory systems, with robotic arms and active vision, were the subject of the European project MindRACES between 2004 and 2007, in which we took part by developing the AKIRA framework.
From NeRF to Gaussian splats
Reconstructing a 3D scene from photographs has gone through two stages in recent years. In 2020 NeRF, with Mildenhall as first author, showed that a neural network can represent a scene as a function returning density and colour for every point and direction, from which volume rendering produces images from viewpoints never photographed, starting from dozens of photos with known camera positions. In 2023 3D Gaussian Splatting by Kerbl and colleagues replaced the network with millions of 3D ellipsoids, each with a position, shape, colour and opacity, that rasterise in real time. This is the representation Marble and Atlas use.
A Gaussian splat describes well how a scene looks and poorly how it behaves: it holds no masses, no friction and no solid surfaces. That is why Marble also exports collision meshes, and why in the real-to-sim-to-real method the hard part is dynamics. It is the same distinction between renderer and simulator, seen from the side of file formats.
What World Labs has built
In less than three years the company has lined up these steps:
- 12 November 2025: Marble, a multimodal model that generates explorable 3D worlds and exports 3D Gaussian splats and collision meshes;
- 21 January 2026: a public World API to generate 3D worlds from text, images and video;
- 21 July 2026: the acquisition of SceniX, a robotics and simulation company that was developing a real-to-sim-to-real engine;
- 1 September 2026: Atlas, the new world model.
Atlas is described as an omni model trained from scratch on text, images, video and 3D, with an autoregressive diffusion transformer architecture. All inputs enter a shared spatial context in which every image has a position in 3D space, and the model generates what comes next while staying consistent with what it has seen. According to World Labs it reconstructs a scene from one to more than a hundred images, faithfully with as few as two or three; it produces video of up to one minute at 1440p; and it returns explicit 3D outputs such as point clouds and Gaussian splats. For robotics the company shows two large environments reconstructed from a phone video, 24 frames each, in which Atlas generates the RGB images and depth that a simulated robot’s cameras would see as it moves.
These are results stated by the company, published as technical posts with demonstration videos.
Real-to-sim-to-real, step by step
Training in simulation and transferring to the real robot is not a new idea. The best-known technique is domain randomization: a simulator is built by hand and its colours, lighting, masses and friction are varied at random until the policy learns to ignore the differences from reality. With an automatic version of this approach OpenAI trained a robot hand to solve a Rubik’s cube in 2019. The limit is the cost of modelling every environment and every object by hand.
The method World Labs presented with SceniX on 28 July starts instead from the capture of a real task and uses a generative model to reconstruct it. In short:
- Capture: the robot, sensors, environment, objects and a few demonstrations of the task are recorded.
- Reconstruction: the task becomes an interactive world that preserves appearance, geometry and dynamics. Where the shape, weight or friction of objects is uncertain, and where cables and packaging deform, the system combines different representations depending on what matters for the task.
- Validation: the same action sequence runs open loop in simulation and in reality, comparing observations, object motion and outcomes.
- Variation: appearance, object configuration, clutter, physical parameters, robot state and camera viewpoint are varied systematically.
- Training: policies learn in simulation, their failure points are searched for and the experience needed to fix them is generated.
- Evaluation: checkpoints are screened in simulation and only the most promising reach the hardware.
Two results are stated. Policies trained with zero real-world data were transferred directly to five different platforms (ALOHA, YAM, RB-Y1, Flexiv and xArm), and several of them ran for an hour without intervention on cables, test tubes and thin objects in clutter. Evaluation was measured on one task, a cube handover between two ALOHA arms: each checkpoint was tested on 2,000 simulated trials and 100 real ones, half inside and half outside the training distribution, and the simulation preserved the ranking between policies and the regions of success and failure.
The criterion World Labs puts in writing applies even to those who do not use its tools: a useful simulation does not need to reproduce real success rates exactly, it needs to lead to the same decisions as reality, namely which policy is better, where it fails and whether an improvement in training carries over to hardware.
The limits are those of any result not yet replicated: the data is the company’s, there is no published independent validation, evaluation is shown on a single task and the tasks are bench-top manipulation, not a complete production line. World Labs’ own taxonomy adds a risk specific to generative simulators: generated geometry can “look correct while containing self-intersections or wrong scale that produce nonsensical physics”.
Why a chipmaker is buying it
AMD justifies the deal with workloads. The press release explains that as AI expands into reasoning, robotics, simulation and physical AI, the demands on compute infrastructure become more diverse, and that World Labs’ expertise will give AMD a more direct view of how workloads evolve, useful in shaping its roadmap. The two companies had been working together since 2025 on training and inference optimisation on AMD GPUs, and Li writes that Lisa Su was an early investor.
The deal fits a visible path. In August 2024 AMD completed the acquisition of Silo AI, a Finnish lab its press release called the largest private AI lab in Europe, for about $665 million; in March 2025 that of ZT Systems, which designs rack-scale systems for data centres. The picture is one of a complete chain: silicon, systems, the open source ROCm software stack, model labs and now a group working on a class of workloads different from language models.
Our reading is that world models have a different compute profile. Within the same pipeline sit high-resolution video generation, 3D reconstruction, Gaussian splat rendering, physics simulation and above all a very large number of parallel trials: 2,000 simulated trials per checkpoint, multiplied by the checkpoints and the variants. World Labs also states that Atlas’s performance improves with training compute and expects the trend to hold, and a curve of that kind is what interests anyone designing accelerators. Competitively, NVIDIA already offers a family of open-weight world models, Cosmos, while according to TechCrunch AMD had so far released only text and video models.
Li writes that she wants to provide “the best open models”. As of today the real-to-sim-to-real engine is explicitly proprietary and Atlas has not been released with open weights: what will be published after closing is one of the things to verify.
Letting a shop floor be captured: what to weigh first
In the real-to-sim-to-real method the input is the capture of a real task in a real place: the robot, the sensors, the environment, the objects, some demonstrations. To reconstruct an environment to navigate, in the Atlas example, 24 frames of a phone video are enough. World Labs describes every reconstructed world as reusable infrastructure: once reconstructed, a task can serve many policies and different robots, over time too. For a manufacturer this means that the capture of a cell or a department becomes an asset with a value of its own, and it is worth treating it as such before handing it to a supplier.
- Who can reuse the reconstructed world. The contract has to say whether the reconstruction and the policies trained on it can serve the supplier’s other customers, and for how long.
- Trade secrets. A reconstruction with geometry and dynamics exposes layout, equipment, cycle times and process parameters. Under EU trade secrets law, as transposed in Italy by article 98 of the Industrial Property Code, protection requires the information to be subject to reasonable steps to keep it secret: governing who captures, what, and where the data ends up is part of those steps.
- The people in the footage. Demonstrations are performed by operators. Video of people at work is personal data, and in Italy the use of cameras that may allow remote monitoring of workers’ activity has to be assessed against article 4 of the Workers’ Statute, which allows it only for organisational and production needs, workplace safety or protection of company assets, and subject to a union agreement or authorisation from the labour inspectorate.
- Machine data. For data generated by connected robots and machinery, the Data Act, applicable since 12 September 2025, provides that the data holder, often the manufacturer, may use non-personal data only on the basis of a contract with the user (article 4(13)).
Then there is a technical issue. Atlas fills the parts of a scene it has not seen with what it considers plausible, and World Labs says so openly: the more images it receives, the less it imagines. In a creative application that is a quality; in a cell where a robot has to grasp a part it is a risk. Capture coverage in the areas where the robot comes into contact with objects, and validation of those contacts, are part of the engineering work just as much as the choice of model.
What to watch
- Approvals and closing of the deal, expected by the end of 2026.
- What changes for current users of Marble and the World API, on which the press release says nothing.
- Whether Atlas or other World Labs models will be released with open weights, as Li says she intends.
- Whether World Labs’ tools will be brought onto AMD’s software stack, and on what terms for users of other vendors’ hardware.
- Independent validation of the real-to-sim-to-real method and results on longer tasks and less controlled environments than bench-top manipulation.
Sources
- AMD, AMD to Acquire World Labs to Advance the Future of AI Compute
- Fei-Fei Li, To Seek a Newer World
- World Labs, About
- World Labs, Atlas: A World Model for Spatial Intelligence
- World Labs, Building Worlds That Train Robots
- World Labs, A Functional Taxonomy of World Models
- World Labs, Marble: A Multimodal World Model
- TechCrunch, AMD will acquire Fei-Fei Li’s World Labs for $8.2 billion
- AMD, AMD Completes Acquisition of Silo AI
- AMD, AMD Completes Acquisition of ZT Systems
- Mildenhall et al., NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (2020)
- Kerbl et al., 3D Gaussian Splatting for Real-Time Radiance Field Rendering (2023)
- OpenAI et al., Solving Rubik’s Cube with a Robot Hand (2019)
- Hans Moravec, Mind Children: The Future of Robot and Human Intelligence, Harvard University Press, 1988
- Regulation (EU) 2023/2854, Data Act
- Italian Industrial Property Code, art. 98
- Italian Workers’ Statute, art. 4
- MindRACES: anticipatory cognitive systems
- A.K.I.R.A.
