A Humanoid Was Asked to Change a Light Bulb. NVIDIA's Flagship AI Tied a Robot Doing Nothing At All.
A new open benchmark called Fiatlux puts a Unitree G1 humanoid through climbing a ladder and replacing an overhead light bulb in one continuous episode. NVIDIA's flagship GR00T N1.7 model scored statistically the same as a policy that does nothing, and even expert human teleoperation could not complete the climbing subtasks, pointing to whole-body control, not AI reasoning, as the real bottleneck.

Put a free-standing A-frame step ladder in front of a humanoid robot. Tell it to climb up, unscrew a dead light bulb from an overhead fixture, carry it down without dropping it, screw in a fresh one, and climb back down. A competent adult does this in under two minutes with no training. A benchmark published on arXiv this week put that exact task in front of the robotics industry's best openly available foundation model, and the model tied, statistically, a policy that issues zero commands at all.
That is the thesis of this column: the humanoid industry keeps measuring progress by how good a demo looks, and a new open benchmark just showed that the real ceiling is not the AI brain everyone is funding. It is the body underneath it. Nobody has closed that gap yet, not the leading vision-language-action model, not a hand-coded controller, and, more uncomfortably, not even a human expert driving the same robot by remote control.
The benchmark is called Fiatlux, built by five researchers, four of them affiliated with the University of Hawaiʻi at Mānoa and one jointly with Purdue University, as a native extension for NVIDIA's Isaac Lab simulator. It targets the Unitree G1, a 1.3-meter, roughly 35-kilogram humanoid that Unitree already sells for a base price of US$16,000. This is not a concept robot. It is hardware you can order today, running a task every building-maintenance worker performs without thinking about it.
The Test Built to Expose a Hidden Gap
Most humanoid benchmarks score locomotion and manipulation separately. HumanoidBench evaluates fifteen whole-body manipulation tasks and twelve locomotion tasks across a 61-degree-of-freedom action space, but each task runs in isolation. RoboCasa and its larger RoboCasa365 extension score kitchen manipulation across more than 150 object categories and 100 evaluation tasks, with no climbing component at all. Neither family asks a robot to carry a fragile object while climbing, because climbing with a payload is a different control problem from walking with one.
Fiatlux fuses the two. One continuous 1,440-second episode, twelve atomic subtasks scored in sequence: move the ladder into position, climb it, remove the old bulb, descend while holding it, carry it to a disposal crate, dispose of it, approach the new bulb, grab it, carry it back to the ladder, climb with it, screw it in, and climb back down. The robot's hands are rated to crush the bulb if they close with more than 50 newtons of force, so the benchmark is also quietly testing whether a machine can be strong enough to climb and gentle enough not to break what it is carrying at the same time.
That is the sentence worth sitting with. A single system has to be strong enough to climb and gentle enough not to break the thing in its hand, at the same moment.
The Number That Should Worry Every Humanoid Investor
NVIDIA's GR00T N1.7 is the obvious model to test against a benchmark like this. It reached general availability in July 2026, is licensed under Apache 2.0, and was trained on roughly 20,000 hours of egocentric human video alongside bimanual and humanoid robot datasets. It is, by any fair description, the most credentialed openly available vision-language-action model in the field right now.
Run zero-shot against Fiatlux's twelve subtasks, GR00T N1.7 scored a weighted 0.0876. A baseline policy that holds every joint at its default position and issues no commands scored 0.0861. A policy that samples random actions scored 0.0569. The gap between the flagship model and doing absolutely nothing is smaller than either score's own standard deviation. The paper's authors state the comparison directly: GR00T N1.7 is not separable from the zero-action baseline.
Zero completions. Not low. Zero. Across all twelve subtasks, at every one of four randomized layout seeds, GR00T N1.7's actual success rate was 0.00. The small positive numbers in its weighted score are partial credit for getting closer to a goal, never for reaching one.
The Human Could Not Climb the Ladder Either
Here is the finding that reframes the whole story. The researchers also recorded expert human teleoperation, an operator wearing a PICO 4 headset, driving the G1's arms through inverse kinematics while a whole-body controller handled the legs. On the eight subtasks that do not involve climbing, teleoperation cleared the success gate almost every time, a weighted score of 0.84. On the four subtasks that do require climbing— ascending the ladder, descending with the bulb, climbing with the fresh bulb, and descending again—there is no recorded teleoperated take that satisfies the benchmark's own success criteria. Not one.
A human, with full intention and direct control of the same hardware, could not reliably get this specific robot up and down an A-frame ladder while carrying something fragile, under the benchmark's own safety bounds. That is not an AI capability problem. It is a hardware and whole-body-control problem that no amount of better language-model reasoning will fix by itself.
The hardest part of changing the bulb was never the bulb. It was the ladder.
Why an ASEAN Buyer Should Care About a Simulation Paper
I read a lot of humanoid vendor pitch decks for this platform, and almost every one of them stays on flat ground: warehouse picking, scripted greetings, guided factory tours. That is not timidity. Fiatlux is a reasonably precise explanation for why. A facilities-maintenance buyer in Manila or Jakarta evaluating a humanoid for exactly the kind of work this benchmark simulates, servicing an overhead fixture that requires a ladder, should treat whole-body climbing with a payload as an unsolved category, not a feature arriving next product cycle.
That does not mean humanoids are vaporware. The eight non-climbing subtasks in Fiatlux—navigating a room, picking up an object, carrying it, and placing it precisely—performed substantially better under teleoperation and improving steadily under autonomous policies elsewhere in the field. The tradeoff is specific: trust a vendor's roadmap for flat-ground inspection and sorting work, and ask hard, benchmark-backed questions about any pitch that implies stairs, ladders, or scaffolding are coming soon.
What Actually Moves This Number
The instinct in this industry is to wait for a bigger model and assume the score rises on its own. Fiatlux argues against that instinct directly. A pretrained, general-purpose vision-language-action model tied a robot that was told to do nothing, which means more parameters and more video hours will not, by themselves, solve a control problem that the researchers' own expert pilot could not solve by hand. What is missing is not language understanding. It is a whole-body controller purpose-built for multi-limb contact transitions, the kind of rung-to-rung weight shift, center-of-mass management, and grip-force discipline that this paper treats as a distinct, unsolved research problem, separate from the vision-language-action stack getting most of the funding headlines.
The next twelve months will bring more humanoid demo videos than the twelve before it, most of them filmed on flat, forgiving floors for good reason. The number worth tracking is not how polished the next video looks. It is whether any released policy, autonomous or teleoperated, posts a nonzero clean success rate on Fiatlux's four climbing subtasks. An Apache-licensed model backed by one of the most valuable companies in computing just tied a robot doing nothing at all. Until that specific number moves, a sales deck showing a humanoid on a ladder is showing a stunt, not a capability.
This analysis reflects publicly available research and product information as of the publication date and should not be read as investment, financial, or professional advice; it is provided for general information purposes only.
Hero image: Unitree G1 humanoid robot, official product photography, Unitree Robotics (unitree.com).












