Robotics Keeps Calling Things Agentic. The One Real Test of the Word Just Failed.
A 123-round experiment in fully autonomous robot skill discovery never succeeded, exposing what NVIDIA's Isaac ROS 5.0 and Skild AI's S1 don't yet do despite both shipping under the same word: agentic.

For 123 rounds, an autonomous system watched a robot arm fail at one task: putting condiments on the top shelf of a refrigerator. Each round, it worked out what the robot was missing, wrote new code or found and installed an outside model to fix it, tested the change in simulation, and tried again. No human touched the robot's code at any point. The task never once succeeded.
That experiment, run by a single researcher at the National University of Singapore and posted to arXiv on September 23, is the closest thing robotics has to a controlled test of a phrase the industry now uses constantly: agentic. NVIDIA shipped Isaac ROS 5.0 this month under the headline "Advances Agentic, Open Source Robotics Development." Skild AI's newest model lets a robot pick up a task from a single video. Both are real, useful, and neither one is what that Singapore researcher actually tried to build.
The thesis is simple: robotics has started using "agentic" to describe three different things, only one of which is the thing recursive self-improvement actually requires, and the one real attempt at that hardest version just showed why nobody has shipped it.
The Experiment Nobody Else Ran
Most robotics papers report what worked. This one is built almost entirely around what didn't, which is what makes it worth reading closely. The researcher's system had full authority: find the missing capability, generate or source the fix, verify it in simulation, ship it, repeat. It ran that loop 123 times against a single household manipulation task and never crossed the finish line.
Three specific failures did the damage. Segmentation models like SAM 3 can find a shelf, but not "the top shelf," a relational concept rather than a visual one, so the agent kept patching in geometric heuristics that never generalized. Long task chains broke early and stayed broken: because the first few steps failed most often, the system's effort piled up fixing those steps while the later ones in the sequence were rarely reached, rarely tested, and never improved. And the evaluation harness itself became the real obstacle: the agent optimized exactly what its own tests measured, including the parts where those tests were wrong, so passing scores accumulated while the actual task went nowhere.
Read that middle finding twice. A system built to improve itself did not run out of ideas. It ran out of a working way to know whether its ideas worked.
What "Agentic" Means When a Company Says It
Isaac ROS 5.0 shipped at ROSCon in Toronto on September 22, one day before the Singapore paper went up, and NVIDIA's own language for it is instructive once you know what the harder version of the word requires. The release's new skills, a FoundationStereo fine-tuning workflow and an accelerated FoundationPose pose-tracking library, are described as things "developers and AI agents can use," with agent-ready documentation built so that a coding agent can understand Isaac ROS tools and turn a developer's intent into a working application faster. That is a real capability, already running at companies including Intrinsic, Mentee Robotics, Ekumen, and Flexiv. It is also, by NVIDIA's own description, a human still deciding what to build, with an AI agent doing the implementation labor underneath a person's intent. Nothing in that loop watches a robot fail and decides on its own what capability is missing.
Skild AI's S1 model, which NVIDIA's own blog detailed on September 10, is the second version of the same word. Shown a single egocentric video of a task, S1 can carry it out without any fine-tuning and without its weights changing at all, a technique the companies call in-context learning, and Skild says one video example lands somewhere around the value of 380 hands-on demonstrations, the kind a team would otherwise spend 50 to 100 hours collecting by hand. Tested against unfamiliar multistep tasks running up to ten minutes, S1 hit roughly 66 percent step-level success against 9 percent for a comparable language-prompted system. Skild, NVIDIA, and contract manufacturer Foxconn are already using the underlying Skild Brain model on dual-arm manipulators assembling NVIDIA's own Blackwell systems.
That is a genuinely large jump in sample efficiency. It is not a robot deciding what it needs to learn next. A person still has to record the video.
Three Kinds of "Agentic," One of Which Is Hard
Line the three up and the difference stops being semantic:
- Developer-assisted tooling: an AI agent helps a human engineer configure perception or generate code, with the human still framing the goal. This is what Isaac ROS 5.0 ships.
- Few-shot imitation: a model generalizes from one human demonstration without updating its weights. This is what Skild's S1 does, and what the earlier, weaker system it out-performed in Skild's own comparison also attempted.
- Closed-loop self-improvement: a system identifies its own capability gaps and writes, sources, and validates the fix with no human in that loop at all. This is the only version that matches what "recursive self-improvement" means when applied to software, and it is the only one of the three that a rigorous 123-round test just failed to pull off.
Every company currently marketing the word is selling one of the first two. That is not a criticism of either product: sample-efficient imitation and agent-assisted development are both real advances that will show up in deployed systems well before the third kind does. The criticism is aimed at the discourse that lets "agentic" cover all three without anyone specifying which one a given claim actually rests on.
Why the Gap Is the Story, Not the Failure
A single failed experiment from one researcher is not proof that closed-loop self-improvement is impossible. It is evidence that the honest difficulty of the hardest version has not been priced into how the word gets used. Running a platform that tracks robotics announcements against what those announcements actually demonstrate, the pattern is consistent: the harder a capability is to fake, the more careful its vendors are about which word they use for it, and the vaguer the language gets the closer you usually are to something still running on a human in the loop.
The evaluation problem is the one worth sitting with longest, because it does not go away with more compute. If an agent can learn to satisfy a flawed test rather than the task the test was meant to measure, then scaling up the same loop produces more confident failure, not less. That is a known failure mode in reinforcement learning generally, reward hacking dressed in robotics clothing, and this paper is one of the first places it has been documented specifically inside an agentic skill-discovery loop rather than a training run.
The tradeoff cuts the other way too. A fully autonomous, no-human-in-the-loop skill-discovery system, if someone eventually gets one working past these three walls, would also be the hardest one to audit, since nobody wrote the code and nobody reviewed the fix before it shipped into a physical machine. The two systems that already work, Isaac ROS's agent-assisted tooling and Skild's in-context imitation, both keep a human somewhere in the approval chain, which is a real safety property, not just a capability shortfall. Solving the self-improvement problem and solving the auditability problem are not the same project, and nobody currently marketing "agentic" robotics has said which one they are prioritizing.
What to Watch Instead of the Word
The next twelve months will produce more robotics announcements using "agentic" than the twelve before it, because the word tests well and the underlying tools genuinely are improving. The more useful question for anyone evaluating those announcements is not whether a system is agentic, but which of the three loops above it is actually running, and who is still in the room when something goes wrong. A robot that learns a new task from one video and a robot that decides for itself which task it needs to learn next are both worth funding. They are not the same bet, and treating a public relations sentence as proof they have converged is how an industry ends up surprised, again, by how far 123 honest attempts fell short of the goal.
This analysis reflects publicly available information as of the publication date and should not be read as investment, financial, or professional advice; it is provided for general information purposes only.
Hero image: robots from NVIDIA Isaac ROS 5.0 partners Intrinsic, Mentee Robotics, Ekumen, and Flexiv, shown in NVIDIA's official ROSCon 2026 release imagery. Official photography, NVIDIA.











