RobotAIGeek

Simate Tops RoboDojo Leaderboard Three Months After Founding

Simate, a three-month-old Beijing startup, has topped the RoboDojo robot-manipulation benchmark with a model developed using an AI-driven automated research pipeline rather than only human engineers.

martti
2 Min. LesezeitPosted: 27. Sept. 2026
Simate Tops RoboDojo Leaderboard Three Months After Founding

A three-month-old Beijing startup called Simate has taken the top score on RoboDojo, a widely watched benchmark for general-purpose robot manipulation, with a model its founder says was designed largely by other AI systems rather than by human researchers alone. Simate-beta posted an average RoboDojo score of 33.95 and a task success rate of 27.96%, ahead of every other entrant on the leaderboard, without any task-specific tuning for the benchmark itself.

Simate is led by Zhang Ying, a former core technical lead on autonomous-driving systems who has described the company's underlying approach as "AI for Physical AI": using automated research systems to design, test and refine robot intelligence models faster than a human engineering team could manage alone. The company has raised multiple funding rounds since its founding three months ago, each valued in the hundreds of millions of yuan, a pace that reflects how much capital is currently chasing physical AI research in China even at the earliest, least-proven stage of a company's life.

Simate is headquartered in Beijing and operates at the earliest stage of what is generally called physical AI or embodied AI: building foundation models that let robots perceive and act in the physical world, as distinct from the large language models that dominate text and image generation. Its core product, Sipai, is the manipulation model family that includes Simate-beta, running on an infrastructure layer the company calls Sinfra; the RoboDojo result is the first public evidence that the approach works well enough to compete with better-funded labs.

An Automated Research Pipeline Built to Out-Iterate Human Teams

What separates Simate from most model releases is not the benchmark score alone but the process the company says produced it. Simate has built three linked systems: SiPAI, a modular framework that makes model architectures legible to automated analysis rather than only to the humans who wrote them; AutoResearch, a system that pulls live context from papers, code repositories, prior experiments and internal data to let an AI agent propose and run its own model changes; and an AI-native infrastructure layer that executes many candidate experiments in parallel simulation before promoting the strongest candidates to real-world robot testing, an approach to closing the sim-to-real gap that parallels the automated data pipelines described in RobotAIGeek's coverage of SimFoundry's training-data automation.

In practice, that means an AI agent inside AutoResearch can modify a model's architecture, adjust its training configuration, launch an experiment, evaluate the result against a target metric and decide what to try next, with a human researcher setting strategic direction rather than writing each experiment by hand. Simate frames this as a spectrum of "Physical RSI," recursive self-improvement applied to robotics, running from a weak form where AI assists with routine experiment execution to a strong form where AI-driven iteration handles most of the research loop while humans retain control over which capabilities to pursue and when a result is trustworthy enough to ship.

The company is now opening AutoResearch to outside researchers, including participants from MIT, Caltech, Tsinghua University and Peking University, turning what began as an internal tool into a shared research platform. That step matters commercially as much as scientifically: if outside labs adopt AutoResearch as their own experimentation layer, Simate's automated-research approach becomes an industry reference point rather than one company's internal claim, in the same way early open-source model-training frameworks shaped how an entire generation of AI labs organised their own work.

What Simate-beta Actually Does Differently

The model itself is built around two capabilities the company calls four-dimensional physical perception and hierarchical temporal memory. The first is meant to give the model a working sense of physical properties, such as an object's weight, friction and center of gravity, rather than treating a scene as a flat image to be pattern-matched. The second is meant to let the model retain context across long, multi-step tasks instead of treating each action as though it starts from a blank state.

Public demonstrations show a robot arm using Simate-beta to complete tasks such as making tea, a sequence that requires remembering which steps have already been completed, adapting grip and force as different vessels and implements are picked up, and recovering cleanly when an intermediate step does not go as planned. Those are precisely the categories where general-purpose manipulation models have historically struggled: short demonstrations of a single skill are common across the industry, while models that hold up across long task sequences with realistic recovery from small errors remain rare. RoboDojo was built to separate those two groups of claims by scoring memory, precision, long-horizon execution and generalisation as distinct capability dimensions rather than a single aggregate task-completion number, which is part of why a leading score on it carries more weight than a leaderboard win on a narrower, single-skill benchmark.

The gap between Simate-beta and the rest of the leaderboard is wide enough to be notable rather than marginal. In the two weeks before Simate's entry, the strongest scores on the board's simulation track were Liber-0 Preview, contributed by LiberAI, at 30.74 average score and a 25.52% success rate, and an entry known on the board as GPT-6-Astra at 28.97 and 22.48%. Two other recent entrants, DeepSeek-Flash and a model labelled GPT-5.5, scored below 3.0 and under 1% success under the board's harder evaluation settings, illustrating how steeply performance drops off once a model is pushed past its training distribution. Simate-beta's 33.95 and 27.96% clears the next-best entry by roughly three full points on the average-score measure, achieved, according to the company, without tuning the model specifically for RoboDojo's task set.

The Autonomous-Driving Talent Migration Behind the Model

Zhang Ying's background is the more structurally interesting part of this story. Autonomous driving produced an entire generation of engineers who spent years solving problems that map closely onto embodied AI: real-time perception under sensor noise, planning under uncertainty, and building systems that must work reliably outside a lab rather than merely in a demo. Zhang's team is described as having built a driving stack comparable to Tesla's Full Self-Driving system through three architectural generations, from map-based to mapless to end-to-end learning, giving the team direct experience with exactly the kind of iterative, data-driven model development that AutoResearch is now trying to automate.

Simate's other core hires reinforce the same pattern. Zhan Fangneng, an assistant professor at the Hong Kong University of Science and Technology who directs a lab focused on world models, and Ji Mazeyu, a researcher previously with a Meta-acquired physical-AI startup focused on humanoid robot control, both moved from adjacent research fields into general-purpose manipulation rather than starting there. That migration is not unique to Simate. As autonomous-driving investment cycles have matured and slowed relative to the newer capital flooding into embodied AI, a steady stream of senior engineers has moved from self-driving companies into humanoid and manipulation robotics, bringing production engineering discipline that pure robotics-research teams often lack, a shift examined in RobotAIGeek's analysis of the world-model era's economics, which is a useful reference point for readers trying to place Simate within a broader wave rather than as an isolated event.

Why a Three-Month-Old Leaderboard Win Should Be Read Cautiously

A RoboDojo score achieved three months after founding is a genuinely fast result, but it is also an early one. Leaderboard scores in fast-moving benchmarks are frequently overtaken within weeks as competitors adjust their own models to the same test suite, and a benchmark win says nothing directly about whether a model can be deployed reliably on a paying customer's factory floor or in a home, where lighting, clutter and object variety are far less controlled than in a benchmark's standardised task set. Simate has not disclosed a commercial deployment, a paying customer or a production timeline; the announcement is a research and recruiting milestone before it is a business one.

The more durable claim to watch is not this specific score but whether AutoResearch actually compounds. If AI-assisted iteration lets Simate ship meaningfully improved models every few months rather than on the multi-year cycles typical of foundation-model development, the company's advantage would come from research velocity rather than from any single architectural idea, and that kind of advantage tends to be harder for competitors to copy quickly. If the automated pipeline instead produces diminishing returns once the easy optimisations are exhausted, a first-mover benchmark score alone will not sustain a company competing against far better-funded labs already building their own automated research tools. The company's promised release of papers and open-source contributions by year end will be the first real test of which pattern is playing out.

For procurement teams and automation buyers watching the embodied AI market rather than participating in its research race, the practical takeaway is patience rather than urgency. A benchmark leader with no disclosed customer, no published safety validation and no manufacturing partner is a signal to monitor, not yet a vendor to shortlist. The more useful question to revisit in three to six months is whether Simate converts research velocity into a model that a systems integrator can actually license and deploy, since that gap, between a leaderboard result and a supportable commercial product, is where most fast-rising physical AI startups over the past two years have stalled.

This analysis synthesizes company statements, technical demonstrations and public reporting on Simate's RoboDojo results as of September 27, 2026. It is for general information purposes only and does not constitute investment or professional advice.

Hero image credit: Simate.