RobotAIGeek

Simate Claims It "Beat GPT-6" at Robotics. Its Own Leaderboard Says 27.96 Percent.

A three-month-old startup's claim to have "beaten GPT-6" on a robotics benchmark went viral unverified this week, while the benchmark's own published numbers show even the leaderboard-topping model still fails most real-world manipulation tasks against a 100 percent human baseline.

martti
4 min readPosted: Sep 27, 2026
Simate Claims It "Beat GPT-6" at Robotics. Its Own Leaderboard Says 27.96 Percent.

A company that did not exist in March topped a robot-manipulation leaderboard this week, and the framing that spread fastest said it had "surpassed GPT-6-Astra and DeepMind" to take the global top spot. No score accompanied that claim. No date, no leaderboard entry, no way to check it against anything. It just travelled, the way a claim like that is built to travel.

Here is what did not travel with it: the benchmark this company topped publishes its own numbers, and those numbers say the field is still barely walking. A human completing the same manipulation tasks scores 100 percent in the real world and 76 percent in simulation. The best model on record before this week's leaderboard update, an open system called Hy-Embodied-0.5-VLA, cleared 8.8 percent in simulation. On tasks that required understanding an unfamiliar spoken instruction rather than a memorized one, the top score across every model tested was 1.67 percent. That is the actual state of the art the "beat GPT-6" claim is standing on. My thesis is simple: embodied AI just re-imported the exact benchmark-hype failure mode that discredited early language-model leaderboards, at the precise moment its own evaluators are proving how far the field still has to go.

The Number Nobody Repeated

The benchmark is called RoboDojo, built by the team behind RoboTwin and published as a formal paper in July, covering 42 simulated tasks and 18 real-world tasks across three physical robot arms, scored across five separate capability dimensions: generalization, memory, precision, long-horizon planning, and open-ended semantic understanding. It is not a marketing leaderboard. It is an attempt to give embodied AI something language models already have, a standardized, published, peer-reviewable difficulty curve.

That curve is brutal by design, and the paper says so directly: models that look competent on a rehearsed task collapse when the instruction changes even slightly, and the collapse is worst on exactly the dimension that matters most for a robot working in a home or a warehouse it was not trained in. A benchmark built to expose that gap is doing its job when the scores come back low. The problem is not the benchmark. The problem is what happens the moment a new number gets bolted onto it.

What Actually Changed This Week

On September 23, a new entrant called Simate-beta was added to RoboDojo's simulation leaderboard with a score of 33.95, or 27.96 percent normalized. Compared to July's best published score of 8.8 percent, that is a real jump, and I want to be careful not to erase it: something in that model's training pipeline is meaningfully better than what existed ten weeks earlier. Simate itself is barely older than the improvement. It was founded in June by a former executive from an autonomous-driving company, an HKUST assistant professor whose lab studies world models, and an engineer who previously helped build a physical-AI startup later acquired by Meta. It has raised an undisclosed sum described only as "hundreds of millions of RMB" (tens of millions of US$ at current exchange rates), across multiple rounds in three months, from investors it has not named.

None of that is disqualifying. Fast money chasing a real technical gain is how this industry has always worked. What should give any reader pause is the next fact: the claim of a leaderboard-topping result reached the public before any independent party had re-run it. One of the more careful accounts of the story said so in plain language, describing the result as self-reported, with no independent re-verification, and noting that Simate's own technical report was being held back for later. A 27.96 percent score, unverified, is being read by the market as "beat GPT-6."

A Tale of Two Leaderboards

Compare that to RoboCasa, a rival benchmark from a separate academic team that has been running long enough to build in the lesson RoboDojo has not yet had to learn. RoboCasa's own submission rules require any closed-source model to grant the benchmark's maintainers private access to the model and its code before a score is posted. Open-source submissions get checked the other way, in public, because the code is sitting right there in the pull request. On that leaderboard, three different verified systems now clear 70 to 80 percent on their task set, one GR00T release paired with a world-model add-on at 72.6 percent, a Xiaomi-built policy at 74.5 percent, another system at 79.2 percent.

Sit with the contrast for a second. Two benchmarks, two research communities, two different pictures of how far embodied AI has come, and the difference is not primarily about whose robots are smarter. It is about task design, five hard dimensions versus one well-rehearsed kitchen, and about verification, a private-access requirement versus a still-informal leaderboard update. A reader who compares RoboDojo's 27.96 percent against RoboCasa's 79.2 percent and concludes anything about relative model quality has learned nothing true. That comparison is not apples to oranges. It is not even produce.

What the Hardware Actually Costs

I keep a mental corrective for hype cycles like this one, built from years of maintaining a taxonomy of what robots are and what they actually do: go find the part of the story nobody is marketing. In this case, it is the arm. RoboDojo's real-world tasks run on, among other platforms, AgileX Robotics' Piper X, a six-axis arm listed on the company's own site at US$1,999 before shipping and tax. The model being celebrated as a GPT-6 killer is, in its physical form, a desktop-priced arm executing tasks it still fails roughly seven times out of ten in simulation and worse in the real world. That is not a knock on the hardware, which is a genuinely capable piece of engineering at that price. It is a corrective to a headline that wants you picturing something closer to a humanoid marching into a factory tomorrow.

What I Would Watch Next

The honest version of this story is not "China leapfrogged the West at robotics." It is "a benchmark built by serious researchers is doing exactly what it was designed to do, and the press cycle around one of its new entries is doing exactly what press cycles around leaderboard claims always do." I watched this same pattern happen to language models for two straight years before evaluators built held-out test sets and third-party audits people actually trusted. Embodied AI is now living through its own version of that same lesson, on a harder problem, with a shorter memory.

What I would watch for next is not whether Simate's number holds up, though I hope someone checks. It is whether RoboDojo's maintainers adopt something like RoboCasa's private-access requirement before the next viral leaderboard claim arrives, and whether the next founder who tops that leaderboard publishes the technical report before the headline instead of after it. Until one of those two things happens, the honest reading of any "beat GPT-6" claim in this space is the same: not verified, not yet, and probably not what it sounds like.

This column reflects the author's own analysis of public benchmark documentation, company statements, and academic publications current as of September 27, 2026. It is for general information purposes only and does not constitute investment, financial, or legal advice.

Hero image: AgileX Robotics' PiPER-X robotic arm, one of the physical platforms used in RoboDojo's real-world evaluation tasks, official product photography. Credit: AgileX Robotics.

EmbodiedAIRoboticsBenchmarksPhysicalAIWorldModelsChinaRoboDojoAgileXRobotics