RobotAIGeek

Wujie Power World Model Tops RoboCasa, Beats GR00T

A one year old Chinese embodied artificial intelligence startup, Wujie Power, released a latent space world model called MWA that reached the top of the Stanford backed RoboCasa GR1 TableTop leaderboard with a 75.2 percent average task success rate, ahead of the NVIDIA GR00T-N1.6 model. The result puts a contrarian technical approach, latent space world models paired with reinforcement learning, at the front of a benchmark that tests how robots handle messy, unpredictable conditions.

A
4 min readPosted: Jun 30, 2026
Wujie Power World Model Tops RoboCasa, Beats GR00T

When a person picks up a cup of water, the brain estimates the weight, predicts how much the surface will slosh, and steers clear of the glass next to it, all in well under a second. That kind of physical intuition is exactly what robots have struggled to learn, and it is the problem a one year old Chinese startup says it has now pushed to the front of a major industry benchmark.

The Capability

On June 29, 2026, Wujie Power, known in Chinese as Wujie Dongli, released what it describes as a latent space world model called MWA, built around a long horizon, bidirectional physical causal chain. On the RoboCasa GR1 TableTop leaderboard, a benchmark jointly initiated by Stanford University and other leading institutions, the company's MWA WALA model recorded a 75.2 percent average task success rate and took the top position, ahead of mainstream models including the NVIDIA GR00T-N1.6. The benchmark is demanding by design. It covers 24 high difficulty tasks across non standard kitchen environments, including long horizon multi step processes and object retrieval in constrained spaces, and it adds randomized lighting, clutter, and changes to object specifications to test how a model generalizes in uncertain conditions.

How the Approach Differs

Most embodied artificial intelligence systems today follow the vision language action route, which lets a robot understand a text instruction but tends to break when lighting shifts or an object moves a few centimeters, because the model is imitating demonstrated trajectories rather than understanding physical cause and effect. Wujie Power has taken a less common path that combines a latent space world model with reinforcement learning. The world model builds the robot's understanding of physical rules and predicts future states, while reinforcement learning converts that understanding into precise action through high frequency trial and reward. The model also works entirely inside a shared latent space rather than predicting pixels, which lets it skip wasted computation on irrelevant background detail and instead extract what the company calls latent actions, the underlying representation of how objects change when they are acted upon. Because latent actions do not depend on manual labeling, the model can train directly on vast amounts of unlabeled internet video.

Why It Matters

The headline result is not only the score but the method behind it. By modeling long horizon causal chains rather than single step predictions, the system reduces the snowball effect in which a small early error compounds across a long task until the action sequence collapses. Wujie Power also built a negative sample data system it calls AnyPhys, which deliberately collects tens of thousands of failure, instability, and near miss examples, the kind of data the wider industry tends to ignore in favor of clean successes. In precision insertion tests, the company reports that task success rates under noisy data improved by up to five times. For buyers evaluating embodied artificial intelligence for real factory and service environments, the relevant signal is robustness under disturbance, which is precisely what the RoboCasa conditions are meant to measure.

The Bigger Signal

A contrarian technical route just produced the leading score on a respected benchmark, which matters because the embodied artificial intelligence field has begun shifting its attention from demonstrations to delivery. The harder question for the sector is no longer whether a robot can perform a scripted task on a stage, but whether it can keep working when conditions change without being retrained for every new scene. A model that genuinely understands gravity, contact, and friction does not need to be taught each scenario one by one, because it can generalize on its own. That is the most difficult path toward general embodied intelligence, and it is also the most fundamental, which is why a benchmark result built on physical causal reasoning carries weight beyond a single leaderboard ranking.

Disclaimer: This article is provided for general information purposes only and does not constitute investment, financial, or commercial advice. Benchmark figures and technical claims are attributed to the cited primary source and the company's stated results.

Image credit: Generated illustrative image, brand-neutral, for editorial use.

Source:

·      QbitAI (量子位), "全球首个:隐空间世界模型,打通长时序双向物理因果链了" (Wujie Power MWA latent space world model, RoboCasa leaderboard result), June 29, 2026. https://www.qbitai.com/2026/06/439891.html

·      Benchmark: RoboCasa GR1 TableTop leaderboard (initiated by Stanford University and partner institutions); MWA WALA developed with the deep reinforcement learning team at the Institute of Automation, Chinese Academy of Sciences.