RobotAIGeek

SimFoundry Automates the Missing Link in Robot Training Data

A joint research team from NVIDIA GEAR, Fei-Fei Li's Stanford lab, and Georgia Tech has introduced SimFoundry, a system that automatically generates interactive, trainable robot simulation environments from a single real world video.

A
2 min readPosted: Jul 6, 2026
SimFoundry Automates the Missing Link in Robot Training Data

NVIDIA GEAR and researchers from Fei Fei Li's Stanford group have released SimFoundry, an automated pipeline that turns a single real world video into a physics ready interactive simulation environment without human modeling. This solves the primary bottleneck in embodied AI: the prohibitive cost of collecting real world robot training data. By generating nearly unlimited digital cousins of physical spaces, it allows models like Vision Language Action systems to train in simulation and transfer to reality zero shot.

What Happened

The paper, titled SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation, was submitted to arXiv on June 26, 2026, by a coalition of researchers from NVIDIA's Generalist Embodied Agent Research lab, Georgia Tech, Stanford University, UT Austin, and the University of Toronto. The author list bridges major institutional powers in AI robotics, including Yuke Zhu, Danfei Xu, Jim Fan, and Fei Fei Li.

Crucially, SimFoundry is not a new robot brain model. It is an automated system that builds the training grounds for robot brains. Historically, training a robot required either expensive physical data collection farms or manually engineered simulation environments. SimFoundry replaces both. It takes an ordinary RGB video of a kitchen or a warehouse and runs it through a three stage pipeline: extraction, generation, and augmentation.

First, it uses foundation models to segment and extract every object in the video. It utilizes deep learning techniques to estimate depth and generate point clouds, which are then parsed by vision language models and segmentation tools like SAM 3. Next, it generates 3D meshes, infers joint structures for things like drawers, and annotates physical properties like mass and friction, exporting a digital twin into the IsaacLab physics engine. Finally, it creates digital cousins by automatically altering object appearances, shifting layouts, and deriving new manipulation tasks. In tests across seven manipulation tasks, policies trained entirely inside SimFoundry data transferred zero shot to real robots with up to 100 percent success rates on unseen objects.

The introduction of digital cousins is a significant leap. By taking a digital twin and modifying it to create object cousins (where appearance changes but function remains), scene cousins (where layouts shift), and task cousins (where new interactions are derived), the system multiplies the value of a single video exponentially. This creates a vast, procedurally generated training curriculum that prepares robots for the chaos of the real world.

Why It Matters

The central pain point in AI robotics today is data scarcity. Large language models scaled because the internet provided trillions of text tokens for free. Embodied AI lacks an internet of physical interaction data. The traditional solution, Sim2Real, required armies of engineers to build synthetic environments, a process that is slow, expensive, and difficult to scale. The emerging alternative, Real2Sim, attempts to scan reality into simulation, but prior tools only solved pieces of the puzzle, either reconstructing a static 3D scene without physics or requiring heavy manual configuration.

SimFoundry closes the loop. It is a fully automated Real2Sim2Real pipeline. A single video of a real workspace now yields an infinite procedural training ground. When Vision Language Action models like pi zero point five were fine tuned on SimFoundry data, their success rate on held out tasks jumped from zero to 29 percent. This means robotics companies can now generate their training data synthetically from basic video sweeps of a target facility, bypassing the need to deploy physical data collection fleets.

For companies building general purpose robots, this pipeline shifts the bottleneck from data acquisition to compute. Instead of waiting months to gather enough teleoperated demonstrations to teach a robot a new task, engineers can record a short video, run it through SimFoundry, and generate thousands of varied training scenarios overnight. This accelerates the development cycle and lowers the barrier to entry for creating specialized robot behaviors.

The Value Perspective

From the value perspective, SimFoundry is a structural deflationary event for the training cost component. Data acquisition is currently the most stubborn expense in developing generalist robot policies. By automating the creation of high fidelity simulation environments, the marginal cost of a new training scenario approaches the compute cost of running the pipeline.

Furthermore, the evaluation correlation is striking. SimFoundry's simulation evaluations predict real world performance with a mean Pearson correlation of 0.911. This means hardware companies can test new software releases in simulation with high confidence that the results will hold on the factory floor, drastically reducing the physical testing cycles that currently delay deployment timelines. This level of predictability allows for continuous integration and continuous deployment pipelines in robotics software, mirroring the rapid iteration cycles seen in traditional software engineering.

The economic implications are profound. If a company can reliably evaluate a new grasping policy in simulation and know with over 90 percent certainty how it will perform on a real robot, the capital expenditure required for physical testing fleets drops significantly. This shifts the economic advantage toward companies that can leverage massive compute resources for simulation, rather than those with the largest physical footprint.

Global Competitive Field and Timelines

The race to solve the simulation data bottleneck splits cleanly along geographic and institutional lines. In the United States, the approach is capital heavy and software driven. NVIDIA leads with Cosmos, a world foundation model platform with over 2 million downloads, and tools like SimFoundry. Fei Fei Li's own startup, World Labs, which raised USD230 million, is commercializing generative 3D worlds through its Marble product. DeepMind is pursuing generative interactive worlds with Genie 3.

In China, the approach leans heavily toward physical data factories and state backed research. Companies like AgiBot and Unitree rely on extensive teleoperation fleets to gather real world interaction data, treating physical data collection as an infrastructure moat. Meanwhile, the Beijing Academy of Artificial Intelligence recently released the Wujie Physis zero point one general world foundation model to push the simulation frontier.

SimFoundry operates on an immediate timeline. As an open research release backed by NVIDIA's research budget, its code and methods are available for labs to adopt now. Because it is modular, meaning it simply swaps in better segmentation or 3D generation models as they are invented, its capabilities will scale automatically. It requires no standalone funding; its value accrues to the broader NVIDIA ecosystem by driving adoption of IsaacLab and Omniverse compute. This open research strategy positions NVIDIA as the foundational layer for AI robotics development globally.

Hard Truth

While zero shot transfer from simulation to reality is the holy grail of robotics, the SimFoundry benchmarks are still bounded laboratory tasks like stacking dishware or storing markers. Industrial deployment requires handling edge cases, sensor noise, and hardware degradation that synthetic data, no matter how varied the digital cousins are, struggles to perfectly emulate. Real world data collection farms will remain necessary for the final mile of reliability. The gap between a simulated kitchen and a chaotic factory floor filled with unpredictable human workers is still vast, and bridging it will require more than just better digital cousins.

Sources

[1] SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation. arXiv:2606.28276. June 26, 2026. https://arxiv.org/abs/2606.28276

[2] NVIDIA GEAR Research. SimFoundry Project Page. https://research.nvidia.com/labs/gear/simfoundry/

[3] QbitAI. Sim2Real is too expensive, Real2Sim provides volume. July 4, 2026. https://hub.baai.ac.cn/view/56079