Two Physical World Models Launched on the Same Day. They Could Not Be More Different.
On June 28 and 29, 2026, two companies on opposite sides of the Pacific released physical world models within hours of each other. Decart, a San Francisco-based Artificial Intelligence (AI) lab, opened Application Programming Interface (API) access to Oasis 3, a real-time generative world model built for autonomous vehicles and robotics.

On June 28, 2026, Decart, a San Francisco-based AI lab, opened Application Programming Interface (API) access to Oasis 3, a world model that generates photorealistic driving environments in real time. On June 29, Wujie Power, a Beijing-based embodied AI startup, published its release of MWA, a latent-space world model that topped the RoboCasa GR1 TableTop benchmark for robotic manipulation with an average task success rate of 75.2%. Both companies used the same language to describe what they had built. Both are trying to solve the same fundamental problem. The gap between their approaches is a precise map of the deepest unresolved argument in physical AI.
The Problem Both Are Trying to Solve
Training a robot or an autonomous vehicle requires enormous quantities of data showing how the physical world behaves. Collecting that data in the real world is slow, expensive, and constrained by the limits of what can be safely and legally recorded. The solution the industry has converged on, at least in principle, is a world model: a system that can generate synthetic environments accurate enough that a control policy trained inside them will transfer to real hardware without falling apart.
The word "accurate" is doing a lot of work in that sentence. Accurate enough for what? For a perception system learning to recognise stop signs in rain, the bar is visual realism. For a robot arm learning to pick up a glass without breaking it, the bar is physical causality. These are different problems, and the two companies have built different machines to solve them.
Decart: The Generative Bet
Decart was founded in 2023 in San Francisco by Dean Leitersdorf and Moshe Shalev. The company's founding thesis was not about robotics. It was about speed: that generative AI could be made fast enough and cheap enough to run interactively in real time. That infrastructure bet, packaged as the Decart Optimization Stack (DOS), became the engine underneath everything the company has built since.
The public understood what Decart was doing in October 2024, when the company posted a demo of Oasis 1, a playable version of Minecraft generated entirely by a neural network with no game engine running underneath it. Every block, every frame, every physics interaction was produced on the fly by a model that had watched enough gameplay footage to learn how the world behaved. The demo went viral. Within weeks, Decart had raised a $21 million seed round, followed by a $32 million Series A at a $500 million valuation. By May 2026, after a $300 million raise that brought Toyota, Adobe, and Nvidia in as strategic investors, the valuation had reached nearly $4 billion.
Oasis 3 is the application of that same generative architecture to physical AI. The model generates multi-camera photorealistic driving environments from a text prompt, producing a continuous, interactive scene that an autonomous vehicle developer can use to test perception systems and edge-case scenarios. The pricing is $0.02 per second of generated environment. Leitersdorf has described the efficiency advantage as more than an order of magnitude cheaper than competing approaches.
The limitation is well understood and openly acknowledged. Oasis 3 does not enforce physical consistency with the precision of a deterministic physics engine. In testing reported by TechCrunch, vehicles passed through other vehicles in some scenarios. Road conditions shifted unpredictably between frames. The model's auto-regressive architecture fills its context window quickly, roughly 8,000 tokens per frame at tens of frames per second, which means its effective memory is limited. As the context window fills, consistency degrades. For applications where a control policy needs to transfer directly from simulation to hardware, this is a real constraint. For applications where the goal is generating diverse visual data or stress-testing rare scenarios, it is less critical.
Decart's bet is that the developer ecosystem will find the applications that matter, and that the infrastructure layer, not the application layer, is where the durable business sits. Toyota's participation in the May 2026 round was not subtle. The world's largest automaker does not write large cheques into AI infrastructure companies unless it sees a direct application to its own development pipeline.
Wujie Power: The Latent-Space Bet
Wujie Power was founded in 2025 in Beijing by veterans of Horizon Robotics and Li Auto. The company raised over $200 million in an angel round announced on June 26, 2026, led by JD.com-affiliated funds with participation from Sequoia China and Linear Capital. It is one of the fastest-capitalised embodied AI startups in China's recent wave, and it has secured nearly $100 million in global orders within its first year of operation.
The company's founding argument is a direct critique of the approach that Decart and most Western AI labs have pursued. Wujie Power contends that end-to-end Vision-Language-Action (VLA) models, which process visual input, language instructions, and action outputs in a single pipeline, have a structural ceiling. When these models are pushed into dynamic, multi-environment real-world settings, they lose the ability to self-predict and self-evolve, because they have no internal representation of how the physical world actually works. They can imitate. They cannot reason about physics.
MWA, the model released on June 29, is Wujie Power's answer to that problem. Rather than generating visual output, MWA operates entirely in latent space. It builds an internal representation of physical causality: a model of how actions cause world states to change, and how observed world states constrain what actions are possible. The architecture uses bidirectional dynamics, with a forward model predicting how an action will change the world and a reverse model compressing observed outcomes back into latent information to stabilise policy learning. The output is not a video or an image. It is a sequence of latent action chunks, multi-step predictions of what the robot should do next, produced continuously without the token-by-token bottleneck that limits auto-regressive video generation.
The training recipe follows what the company describes as a "learn first, then act" logic. In the pre-training phase, MWA is trained on large-scale unlabeled video data from the internet, building a general understanding of physical cause and effect without requiring task-specific annotation. In the reinforcement learning phase, the model fine-tunes its action policy using a proprietary data system called AnyPhys for Reinforcement Learning (RL), which Wujie Power describes as the industry's first negative-sample core data system for embodied RL. AnyPhys builds hard-negative samples, physically challenging edge cases, and boundary-failure scenarios to train the model to recognise and recover from the situations where it is most likely to fail. The company claims that in online precision insertion tasks, success rates under noisy conditions improved by as much as 15 times using this approach.
On the RoboCasa GR1 TableTop benchmark, conducted jointly with the Chinese Academy of Sciences (CAS) Institute of Automation, MWA achieved an average task success rate of 75.2%, ahead of ACE-EGO-0 (72.8%), DIAL (70.2%), ABot-M0 (58.3%), and GROOT-N1.6 (47.2%). The benchmark covers robotic manipulation in irregular kitchen environments with variable lighting and object interference, one of the more demanding generalisation tests currently in wide use. The margin over second place was 2.4 percentage points.
Why the Architectures Diverge
The two companies are not building different versions of the same thing. They are answering different questions about what a physical world model is for.
Decart's architecture is optimised for generation speed and visual coverage. Its value is in producing diverse, photorealistic environments quickly and cheaply, so that a perception system or a scenario-testing pipeline can be fed with data it could not collect in the real world. The physical consistency limitation is a known trade-off, not an oversight. For autonomous vehicle developers who need millions of rare-scenario images for perception training, visual diversity at low cost may matter more than perfect physics.
Wujie Power's architecture is optimised for physical reasoning and long-horizon action stability. Its value is in building a robot that understands why actions produce outcomes, not just what sequence of actions tends to work in a given context. The latent-space approach avoids the pixel-level overhead of video generation entirely, which is why the model can produce continuous multi-step action predictions without the context-window degradation that limits Oasis 3. The trade-off is that MWA produces no visual output at all. It cannot generate a training environment. It generates a policy.
These are complementary gaps, not competing products. A robotics team building a full training pipeline might plausibly use a generative world model like Oasis 3 to produce diverse visual training data for perception, and a latent-space model like MWA to train the action policy that operates on top of that perception. The industry has not yet converged on a standard architecture that does both well, which is precisely why two well-capitalised companies released two fundamentally different approaches within hours of each other.
The Deeper Argument
The simultaneous release is not a coincidence of timing. It is a reflection of where the physical AI field currently sits. The industry has agreed that world models matter. It has not agreed on what a world model is supposed to do.
The video-generation school, represented by Decart, Wayve's GAIA, Google DeepMind's Genie 3, and World Labs, argues that a world model should produce rich, interactive environments that look and behave like the real world. The latent-space school, represented by Wujie Power and a growing number of Chinese research groups, argues that a world model should capture physical causality in a compressed internal representation, without the computational overhead of pixel-level generation. The benchmark results from RoboCasa suggest the latent-space approach currently has an edge on manipulation tasks. The autonomous vehicle industry's continued investment in generative approaches suggests the visual coverage advantage matters more for perception-heavy applications.
Both schools are right about something. The question is which trade-off matters more for the specific application being built, and whether a future architecture will eventually collapse the distinction.
The physical AI industry spent years debating whether world models were necessary at all. That debate is over. The new debate is about architecture, and it is more technically substantive than the previous one. Decart's Oasis 3 and Wujie Power's MWA are not competing for the same customer. They are competing for the same thesis: that their approach to representing the physical world will become the foundation on which the next generation of robots and autonomous systems is built. The fact that two companies with $4 billion and $200 million in backing respectively chose the same week to stake that claim publicly is the clearest signal yet that the architectural question is no longer academic.
Sources:
• Decart official announcement, June 28, 2026: https://decart.ai/articles/oasis-3
• TechCrunch exclusive, Rebecca Bellan, June 10, 2026: https://techcrunch.com/2026/06/10/decarts-new-world-model-can-simulate-hours-of-photorealistic-driving-with-some-caveats/
• Wujie Power (Anyverse Dynamics) WeChat Official Account, June 29, 2026, 10:02: 登顶具身智能权威榜单!无界动力发布 MWA™ 隐空间世界模型
• RoboCasa GR1 TableTop benchmark, conducted by Wujie Power x Chinese Academy of Sciences Institute of Automation (中科院自动化所)
• TechCrunch, Decart Oasis 1 original launch, October 31, 2024: https://techcrunch.com/2024/10/31/decarts-ai-simulates-a-real-time-playable-version-of-minecraft/
• SiliconAngle, Decart $32M Series A, December 19, 2024: https://siliconangle.com/2024/12/19/ai-world-model-startup-decart-reels-32m/
• Preqin / Exa funding profile, Decart $300M May 2026 round: https://exa.ai/websets/directory/decart-funding












