Beijing Games Work Scenario Events Are the Real Market Map
Behind the sprint-record headlines, the World Humanoid Robot Games ran roughly twenty work-scenario events across nine sectors: factory, hotel, hospital, retail, home, and emergency response. These organiser-defined tasks are the closest thing embodied AI has to an independent benchmark, but the Games still don't publish the failure and intervention data that would make results comparable across vendors.

The Medal Table Is the Marketing. The Work Events Are the Market Map.
While international coverage of the World Humanoid Robot Games in Beijing concentrated on a sprint record, roughly twenty of the competition's events had nothing to do with sport. Robots were sorting library books, making hotel beds, assembling parts, running emergency-response drills, and picking beans off a surface to test fine motor control.
That second programme is the one an industrial buyer should be reading. The Games, which opened on Saturday, August 22 at the National Speed Skating Oval with 2,056 robots from 666 teams across 16 countries, are structured as roughly thirty competitive sports alongside about twenty work-scenario events, and the organisers added fourteen new scenario contests this year spanning nine sectors including households, hotels, and logistics.
What the Scenario Events Actually Test
The settings are deliberately ordinary: factory, hotel, home, hospital, retail, and emergency response. Tasks reported across them include industrial assembly, housekeeping, library book sorting, materials handling, and rescue drills.
Each of those isolates a capability that a sprint does not. Library sorting is a test of reading a label, locating a position in an ordered sequence, and inserting an object into a tight gap without disturbing its neighbours, which is the same problem as warehouse putaway with better lighting. Hotel housekeeping is long-horizon task sequencing in a space where objects move between runs. Emergency response tests behaviour when the environment is degraded and the correct action is not the trained one. Picking beans is a dexterity floor test, and it is the kind of task that separates a hand with real force control from a hand that photographs well.
The events also test something no benchmark suite captures, which is whether a machine can be handed a task specification by a third party rather than choosing its own. In a scenario contest the organisers define the job. The company cannot quietly select the subset of work its robot happens to be good at, which is the standard failure mode of vendor-run demonstrations.
The Benchmark Gap These Events Are Filling
Robotics has no equivalent of the standardised evaluations that shaped other parts of computing. Language models are compared on public test sets that anyone can rerun. Autonomous vehicles report disengagement figures to regulators in some jurisdictions. Industrial robot arms have decades of ISO performance standards covering pose accuracy and repeatability.
Embodied AI has almost none of that. The academic benchmarks that exist, mostly simulation-derived manipulation suites, are useful for research and close to meaningless as procurement evidence, because a score generated in a simulator says little about a specific building with its own lighting, floor surface, and traffic. The result is a market where every vendor publishes its own numbers, generated on its own tasks, with no independent party able to reproduce them.
A physical competition with organiser-defined tasks, a fixed venue, and simultaneous entrants is the closest thing the sector currently has to a controlled comparison. That is a low bar, and it is still the highest bar available.
New Events Worth Noting
This year's additions include tug-of-war and weightlifting, alongside the traditional Chinese sports Taijiquan and Touhu, a pitch-pot arrow-throwing game.
Tug-of-war and weightlifting sound like novelty until you consider what they measure. Both are tests of sustained high-torque output and, more importantly, of what happens to balance and structure when a machine is loaded near its limit against an opposing force it does not control. That is a closer analogue to industrial work, where a robot may be pushing a heavy cart or resisting an unexpected load, than any speed event. Taijiquan is the opposite test, rewarding slow, continuous, precisely controlled whole-body motion, which is unusually hard for control systems tuned to hit discrete waypoints.
Table tennis, contested during the Games, is a perception-latency test in disguise: tracking a small object moving quickly, predicting its trajectory, and committing to a swing before the ball arrives. A team-based football event adds multi-agent coordination, where robots must model teammates and opponents rather than a static environment.
Why This Programme Exists
Beijing has spent two years pushing embodied-AI companies toward public demonstration rather than closed benchmarking, and the scenario events are the clearest expression of that policy. They function as a standardised, third-party evaluation harness at a moment when the industry has no agreed benchmark, no independent certification body for task performance, and a strong commercial incentive to publish only favourable results.
There is an industrial-policy logic on top of that. A national competition with sector-specific categories effectively publishes a list of the applications the state considers priorities, and the companies that perform well in a category acquire a credential that provincial procurement officers can point to. For a domestic manufacturer, a strong showing in the hospital or logistics track is worth more than a medal in the sprint.
The Sectors Worth Watching in the Results
Not all nine scenario sectors carry equal commercial weight, and the ones drawing the least attention are often the ones nearest to real revenue.
Logistics and factory tasks are the most immediately bankable, because the buyers already run automation programmes, already measure throughput, and already accept a capital-expenditure case with a payback period. A robot that performs credibly in an assembly or materials-handling category is talking to customers with budget and precedent.
Hotel and retail tasks sit one step behind: lower stakes if a task fails, but far more variation between sites and much thinner margins to fund the equipment. Hospital work is the most demanding of the group, since a failure carries clinical consequence and procurement runs through committees with a regulatory instinct, but it is also where labour shortages are most acute in ageing economies.
Household tasks, which attract the most consumer interest, remain the furthest from deployment. Homes are the least standardised environments a robot can encounter, and the price ceiling a household will accept is well below what any of these platforms currently costs to build.
The Data the Games Still Do Not Publish
The programme's weakness is the same one that undermines the athletic events. Results are reported as placings and times rather than as attempt counts, intervention counts, and failure modes.
For a scenario event, those omissions matter more than they do on a track. A robot that completes a hotel housekeeping sequence in eight minutes with two human interventions is not comparable to one that takes twelve minutes unaided, yet a results table showing only completion time ranks the first ahead of the second. The metrics that decide a deployment, mean time between interventions, recovery behaviour after a failure, and performance degradation across consecutive runs, are exactly the ones a placing hides.
Publishing them would cost the organisers little and would convert the scenario programme from a showcase into the industry benchmark it is positioned as. It would also be uncomfortable for participants, which is probably why it has not happened yet.
How to Read a Result From These Events
For anyone evaluating vendors, three filters make the scenario results usable.
Ask which category a vendor competed in, not just whether it placed. A company that entered the retail track and finished mid-field has told you more about its readiness for a shop floor than one that won a sprint. Category selection is itself a disclosure.
Ask what the robot did when it failed, since every entrant failed at something across a multi-day programme. Vendors who can describe their failure modes precisely have usually instrumented their systems properly, and instrumentation is what makes a fleet supportable after the sale.
Ask how the task was specified before the run. An event where entrants received the task list weeks in advance is a preparation contest; one where the specification arrives on the day is a generalisation contest, and only the second tells you how a machine behaves on work it has not rehearsed. The organisers do not currently make that distinction explicit in published results, so it is worth asking a vendor directly which of the two its result came from.
Ask whether the same platform competed across multiple scenario categories. A machine that handled both a factory assembly task and a hotel service task is evidence of the general-purpose control stack that most of these companies claim to have built. A different specialised platform per category is evidence of the opposite, however good each one looks individually.
What Comes Next
The Games have grown roughly fourfold in participating robots in a single year, and the scenario programme has grown faster than the sports programme. That trajectory suggests the organisers understand where the durable value is, even if the broadcast schedule and the international coverage have not caught up.
The sprint record will be beaten again, probably next year, and it will generate another round of headlines that compare machines to athletes under rules the machines do not follow. Meanwhile, somewhere in the same building, a robot will fail to fold a towel or misfile a book, and that failure will tell a procurement committee considerably more than nine seconds ever will.
The organisers have built the right competition. What remains is to publish the right numbers.
Image: CGTN, cropped.
Disclaimer: This article is for general information purposes only and does not constitute investment, legal, or procurement advice. Readers should verify details with primary sources before making business decisions.











