RobotAIGeek

Robotics Doesn't Have a Data Problem. It Has an Interface Problem.

Spirit AI's co-founder has set a mid-2027 deadline for a 'GPT-3.0 moment' in robot brains. The same week, a Columbia University lab, Hyundai's own RAI Institute, and an underwater-robotics team all published research on how a human should operate a robot's body, not how the robot should decide on its own. That gap, not the funding round, is the real roadmap.

martti
4 min readPosted: Sep 19, 2026
Robotics Doesn't Have a Data Problem. It Has an Interface Problem.

Between September 15 and 18, three things happened that only make sense read together. Spirit AI's co-founder told the press that robot brains would hit a "GPT-3.0 moment" by the middle of 2027. Hyundai's global head of research and development stood in front of a room of engineers in San Jose and admitted, on the record, that Chinese competitors are "moving far faster" than his own company. And a bimanual robot built by Hyundai's own research arm, the RAI Institute, published a paper whose marquee achievement was catching a baseball.

None of those three events is really about a robot's intelligence. All three are about what still has to be built before intelligence has anything to work with.

The industry keeps talking about a data problem. The papers coming out this month describe an interface problem, and it is a different, earlier, and more stubborn thing to solve.

What Spirit AI Is Actually Disclosing

Spirit AI's pitch is specific enough to respect. Gao Yang, the company's co-founder and chief scientist, calls the robot brain "indeed the weakest link in the complete robotics stack" and has attached a real deadline to closing that gap: a "GPT-3.0 milestone" by mid-2027, at which point he expects a robot to take an open-ended spoken instruction and complete an unfamiliar physical task. Spirit AI has raised more than US$670 million since its January 2024 founding, reaching a valuation near 20 billion yuan (about US$2.9 billion), and its wheeled Moz1 humanoids are working structured shifts inside CATL's battery plant and JD.com's logistics floor today, not in a lab.

The number worth sitting with is not the 2027 deadline. It is the roughly 1,000 contractors, alongside 300 full-time staff, that Spirit AI runs nationwide to operate, correct, and label its robots' physical-interaction data. For tens of deployed Moz1 units, that is a human data-operations team an order of magnitude larger than the robot fleet it supports. Gao's own framing concedes the point: hardware has become commoditized across China's humanoid supply chain, and the scarce input is the labeled interaction data a model needs, data that still requires a person watching, correcting, and re-recording almost everything a robot attempts.

That is not a criticism of Spirit AI, which is more candid about this ratio than most of its competitors. It is a fact about where 2026's humanoid intelligence actually comes from: not from a robot generalizing on its own, but from a standing human operations team that has not yet gotten smaller as the robot fleet has grown.

The Papers Nobody's Funding Round Is Built Around

That is where this month's research agenda becomes the more honest disclosure. A Columbia University lab led by Matei Ciocarlie published DITTO on September 16, a teleoperation interface, not a robot brain. The paper states the actual unsolved problem plainly: existing data-collection methods force a trade-off where "teleoperation ensures deployment consistency but lacks force feedback, while handheld (in-the-wild) systems provide natural force transparency but introduce a visual embodiment gap at deployment." DITTO pairs a seven-degree-of-freedom robotic hand with a motorized exoskeleton built to the operator's own hand geometry, aiming to collect one kind of data that captures both.

Two days earlier, engineers publishing under the RAI Institute, the Hyundai-funded research organization Marc Raibert founded in 2022 with more than US$400 million from Hyundai and Boston Dynamics, released AthenaZero. Its headline achievement is not autonomous reasoning. It is inertia: a 22-actuator, human-mass bimanual body built through quasi-direct-drive actuation to move as fast and as forgivingly as a human arm, about a tenth the effective mass of a conventional industrial manipulator. The robot demonstrates that capability by playing catch. It has thrown a ball at roughly 69 miles per hour, which the research team says is the fastest throw yet recorded from an anthropomorphic robot, caught one at up to 41 miles per hour from 24 feet away, and made bat contact on 27 of 33 pitches, a mechanical feat, not a cognitive one, and the researchers frame it that way themselves.

A third paper posted September 18, ULOHA, extends the same pattern into a corner of robotics with almost no commercial funding pressure at all: underwater manipulation. Its authors note that bimanual imitation learning "has been studied largely in air," and their contribution is custom leader-follower teleoperation hardware built to bring that same human-in-the-loop control underwater. Three papers, three unrelated labs, one shared finding: as of this week, the frontier of published robotics research is still building better ways for a person to operate a dexterous robot, not better ways for the robot to operate itself.

Why Teleoperation Was Supposed to Be a Phase, Not a Destination

The standard industry narrative treats teleoperation as scaffolding, a temporary way to harvest training data en route to genuine autonomy, the same logic behind Figure AI's decision to commit more than US$1 billion to a crowdsourced data platform that has already logged 16 million contributor videos. That framing assumes the scaffolding itself is a solved problem, something companies deploy while they work on the harder cognitive layer above it.

DITTO's own stated trade-off says otherwise. Three-plus years into the current humanoid funding boom, a Columbia lab is still publishing fundamental research on how to capture force-rich manipulation data from a human hand without either sacrificing deployment fidelity or introducing a visual mismatch between the collection rig and the deployed robot. That is not a team polishing an already-working assistive layer. That is a team still defining what the assistive layer should look like.

A "GPT-3.0 moment" forecast that skips past an unsolved interface problem is not forecasting through the hard part. It is forecasting past it.

Text scaled the way it did because the internet had already generated trillions of tokens as an unplanned byproduct of ordinary human communication, and every one of those tokens shared the same representation. Physical interaction data has no equivalent byproduct and no shared representation across robot bodies, sensor suites, and grippers. Every serious lab still has to build the tool that captures it, one exoskeleton, one leader-follower rig, one low-inertia arm at a time, and that construction work is happening in public, dated, and citable, right now.

What This Means for Anyone Actually Buying One

For a procurement team evaluating a humanoid pilot, the diligence question this raises is more useful than asking a vendor for a success-rate figure. Ask instead what fraction of the vendor's training data came from a teleoperation rig purpose-built for their specific hardware, versus a general-purpose collection tool borrowed from elsewhere, and how recently that rig was redesigned. A vendor still iterating on its own data-capture hardware in 2026 is a vendor whose model is still catching up to a moving target, not one converging on a stable one. That is not disqualifying. Spirit AI's own 90 percent structured-task figure, measured across its CATL and JD.com deployments combined, is real, measured performance, not a demo. But it is evidence for a slower, narrower kind of progress than a 2027 deadline implies, and buyers pricing multi-year contracts around that deadline are pricing a bet on an interface layer that the field's own current publication record says is still being invented.

The RAI Institute's baseball robot is, in that light, a more honest data point than any funding announcement this month. It says, plainly, that the hardest solved problem right now is getting a robot's body to move like a person's, quickly, safely, and forgivingly. Teaching that body to decide what to do on its own is the next problem, and the papers dated this week have not started solving it yet. That gap between the two is not a rounding error on the way to 2027. It is the actual roadmap, and it runs through a lab bench before it runs through a factory floor.

This analysis draws on public research publications, company statements, and conference remarks from Spirit AI, Columbia University's robotics researchers, the RAI Institute, Hyundai Motor Group, and Figure AI regarding their respective September 2026 disclosures. It is for general information purposes only and does not constitute investment, financial, or procurement advice.

Hero image credit: RAI Institute, from its own published research on the AthenaZero bimanual robot.

RoboticsHumanoidRobotsPhysicalAIResearchTeleoperationSpiritAIHyundai