RobotAIGeek

Embodied AI Doesn't Have Standards Yet. The Committee Writing Them Meets in Hangzhou This Week.

A UN telecom body's Focus Group on Embodied AI, chaired by China's CAICT with Huawei chairing its parent committee, holds its third plenary in Hangzhou this week to draft the first multilateral terminology and benchmarking standards for embodied AI robots, a process the English-language robotics press has barely covered and in which no major Western humanoid company appears on the documented participant list.

martti
4 min readPosted: Oct 11, 2026
Embodied AI Doesn't Have Standards Yet. The Committee Writing Them Meets in Hangzhou This Week.

Nine working groups, three plenaries, and one keynote from a Chinese government minister later, the body now drafting embodied AI's first multilateral technical standards holds its third meeting this week, in Hangzhou, from October 13 through 16. If that sentence is the first time you have heard of it, you are not alone among people who cover this industry for a living. The committee is the International Telecommunication Union's Focus Group on Embodied AI for Multimedia Technologies, FG-EAI for short, and it is doing the one kind of work that tends to decide a technology's winners long after the product headlines fade: writing down what the words mean.

My thesis is simple, and I think under-covered. The robotics press treats embodied AI as a product race, one humanoid demo against the next. The rulebook— the terminology, the benchmarks, the interoperability specs that will eventually turn up in an enterprise procurement document or a national regulator's reference standard— is being drafted in a venue almost none of that press has mentioned. And the people doing the drafting right now are disproportionately Chinese state research institutions, not through any takeover, but because they are the ones who keep showing up to the unglamorous meetings that actually write the text.

A Chair's Seat, Earned by Attendance

ITU-T Study Group 21 created FG-EAI on February 19, 2026, to study how embodied AI can support what the group calls human-centric multimedia applications: machines that perceive the physical world, learn from trial and error, and work alongside people. Its chair is Yuntao Wang of the China Academy of Information and Communications Technology, a research institute that sits under China's Ministry of Industry and Information Technology and has done this kind of standards-shepherding work before, including on 5G certification. Its two vice-chairs are Shin-Gak Kang of Korea's ETRI and Avinash Agarwal of India's Department of Telecommunications. One level up, the parent study group itself is chaired by Noah Luo, a longtime Huawei standards executive.

None of this is a conspiracy. ITU-T leadership seats are not assigned by decree; they go to the delegates and institutions that turn up consistently to do the actual committee work, meeting after meeting, draft after draft. By that measure, China's state research apparatus is currently winning on attendance, and winning on attendance in a standards body is how previous technology cycles actually got decided.

Nobody voted on which mobile network architecture would win. Committees wrote it into the spec first, and the market followed.

What the Committee Is Actually Writing

FG-EAI's work is split across nine working groups, and the two that matter most are the two nobody outside a standards body finds interesting: terminology and benchmarking. Working Group 1 includes a task group dedicated purely to agreeing on terminology, the literal vocabulary the rest of the field will be required to use. Working Group 8 is building the evaluation and benchmarking framework that regulators and buyers will eventually point to when they ask whether a robot's safety or performance claim is real.

Benchmarks are never neutral.

Whoever's reference architecture becomes the worked example in a published ITU technical report has, in effect, written the default assumptions everyone downstream inherits. The other seven groups, covering multi-modal data, models, systems integration, connectivity, vertical industry applications, human alignment, and a dedicated open-source ecosystem track, read like a table of contents for the humanoid-robot industry's current anxieties: whose data trains the models, whose interfaces the hardware speaks, and whose benchmark decides a robot is safe enough to work near a person.

The Room Is More Global Than the Chair List Suggests

It would be too simple to call this a Chinese operation and stop there, and the evidence does not support that. FG-EAI's July 8 launch workshop in Geneva, the session that formally kicked off the group's public work, drew a genuinely international lineup: Andres Marafioti of Hugging Face, Andrea Cavallaro of EPFL, Wuqiang Yang of the University of Manchester, Sebastian Hallensleben, who chairs the European standards body CEN-CENELEC's own AI committee, and Marija Jankovic, representing ETSI's technical body for conformance testing. The session opened with remarks from China's Minister of Industry and Information Technology, Lecheng Li, and listed a Unitree G1 humanoid as its session robot.

That lineup is the clearest documented snapshot of who was actually in the room, and it shows a real multi-stakeholder process, not a closed one. What it does not show, on the participant list I can verify, is a single named representative from Figure, 1X Technologies, Tesla, Boston Dynamics, or NVIDIA. That is not proof those companies stay away from ITU work entirely; it is what the documented roster from the group's own launch event actually contains. For an industry whose best-funded Western names are usually first in line for a stage, that absence is itself worth noting.

Why Non-Binding Rarely Stays That Way

The honest caveat matters here: ITU-T Recommendations are not law. Plenty of them gather dust, cited by nobody, enforced by no one, because no national regulator ever picked them up. That is the real risk in overstating this story. A standards process this early can still end up mattering to nobody, if the industry simply ignores it and keeps shipping proprietary stacks, which is exactly what most humanoid vendors are doing today.

But the pattern from telecom's last few standards fights, GSM beating out rival network architectures, H.264 outcompeting open-source video codecs for the format embedded in nearly every device, USB-C eventually displacing a decade of proprietary connectors, is that non-binding becomes binding the moment enough national regulators and large buyers start citing the same document in their own procurement rules. By the time that happens, the vocabulary fight is already over, and the companies that were not in the room when the terms got written spend years working around definitions they had no hand in setting.

A standard with someone else's vocabulary already baked in does not need to win the market. It only needs everyone else to stop arguing about what the words mean.

What I'm Watching From Here

Running a platform that tracks company and robot taxonomy by hand teaches a specific kind of paranoia about definitions: a large share of the disputes in classifying a robot come down to vendors using the same word to mean different things, on purpose or not. That is precisely the terrain FG-EAI's terminology and benchmarking groups are fighting over right now, mostly out of sight, while the trade press covers the next demo video.

From an ASEAN seat, where most of the hardware I track arrives built elsewhere and certified against standards written elsewhere, this is not an abstract committee fight. It is a preview of which country's assumptions will show up in the compliance paperwork a regional buyer has to sign years from now. I will be watching two things out of this week's Hangzhou meeting: whether Working Group 8 publishes a benchmarking draft that any government cites before 2027, and whether a single Western humanoid company's name shows up on an FG-EAI working-group roster by the group's next plenary. If neither happens, the vocabulary gets written anyway. It just gets written by whoever stayed in the room.

This analysis draws on the International Telecommunication Union's own published Focus Group on Embodied AI for Multimedia Technologies pages, covering its establishment date, leadership, work programme, and working-group structure; the ITU AI for Good platform's listing of named speakers and agenda for the group's July 8, 2026 launch workshop in Geneva; and public profiles of the China Academy of Information and Communications Technology's role under China's Ministry of Industry and Information Technology. It is for general information purposes only and does not constitute investment, financial, or professional advice.

Hero image credit: Unitree, official H2 humanoid product photography.

RoboticsEmbodiedAIStandardsITUChinaUnitreeGeopolitics