RobotAIGeek

Tally Passed Robotics' Toughest Safety Test. Nobody Checked If It's Actually AI.

Robotics built three real oversight tracks in 2026: UL's safety certification, ISO's robot vocabulary standard, and the EU AI Act's disclosure rule. None of them verifies whether a robot's AI claim is actually true. Tally's UL 3300 certification, NVIDIA's Jetson Orin Nano 2 launch, and Figure AI's voluntary data disclosure show why buyers, not regulators, are currently the only ones checking.

martti
4 min readPosted: Aug 26, 2026
Tally Passed Robotics' Toughest Safety Test. Nobody Checked If It's Actually AI.

In March, a wheeled shelf-scanning robot named Tally became the first public-facing robot to earn UL Solutions' global safety certification under ANSI/UL 3300. The test bench checked whether Tally could catch fire, whether it could shock a shopper, whether it would stop before clipping a grocery cart in a crowded aisle. It did not open the control loop and check what was actually deciding when to stop.

That gap is the whole story of robotics regulation in 2026.

This has been the year the industry finally built serious machinery around robots: safety certification bodies signed off, a vocabulary standard got its third edition, and the world's most ambitious AI law started requiring disclosure. Stack all three tracks together and something is missing. Every one of them certifies something adjacent to the claim printed on the box. None of them certifies the claim itself: that the thing is actually running AI.

The safety track checks the body, not the mind

Simbe Robotics' Tally certification is real progress, and worth taking seriously on its own terms. ANSI/UL 3300, the Standard for Service, Communication, Information, Education and Entertainment Robots, evaluates fire and electric-shock hazards and safe autonomous mobility in spaces where a robot shares floor space with people who do not work for the company that deployed it. On December 31, 2025, the US Occupational Safety and Health Administration added UL 3300 to its Nationally Recognized Testing Laboratory list, which means the standard now carries formal weight in US commercial and enterprise deployments, not just a vendor's marketing page.

None of that testing asks whether Tally's shelf-scanning decisions come from a trained model or a lookup table. It could not, because that is not what UL 3300 was built to measure. A robot that runs a scripted route and a robot running a full perception stack can both pass the same fire-and-shock bench, provided the motor controller behaves the same way in both cases.

The vocabulary track defines the shape, not the intelligence

ISO 8373, the vocabulary standard the International Organization for Standardization maintains for robotics, was revised again in 2021 and now defines terms like robot, autonomy, industrial robot, service robot and medical robot. The International Federation of Robotics and the United Nations Economic Commission for Europe use a parallel classification, splitting the field into industrial robots and service robots, then splitting service robots again into professional and personal use.

Read closely, both frameworks describe what a robot is for and who operates it. Neither one certifies what is running its decisions. A robot classified as a "professional service robot" under IFR's taxonomy has told you its commercial context. It has told you nothing about whether its navigation stack is a neural policy or an if-then tree written five years ago and never updated.

The disclosure track asks for a sentence, not a proof

The European Union's AI Act carries the most consequential transparency rule any jurisdiction has passed for robots that talk to people. Article 50's transparency obligations took effect on August 2, 2026, and they require providers of AI systems that interact directly with individuals to make that interaction clear, unless it is already obvious, at the latest by the first contact. Machine-to-machine interaction, an assembly robot talking to another assembly robot on a closed line, is explicitly out of scope.

Read the obligation for what it actually demands. A company has to tell a person they are interacting with an AI system. Nothing in Article 50 requires the company to prove that what is on the other end of that interaction is, in fact, artificial intelligence rather than a scripted menu tree wearing a chat interface. Compliance is a sentence disclosed at first contact. It was never designed to be an audit of what generated the sentence.

Put the three tracks next to each other and the pattern holds across all of them. Safety certification checks the hardware. Vocabulary standards check the category. Disclosure law checks whether you were told. The one thing none of them checks is the thing every vendor's pitch deck leads with.

NVIDIA just quietly conceded the entry tier has been faking it

On Tuesday, NVIDIA announced the Jetson Orin Nano 2, a robotics compute module rated at 78 trillion operations per second while drawing 40 percent less power than its predecessor in the same form factor. Deepu Talla, NVIDIA's vice president of robotics and edge AI, said the module "puts that breakthrough within reach of millions of developers, delivering the performance and energy efficiency needed for real-time reasoning at the edge."

Read that claim the way a procurement engineer should. If real-time reasoning at the edge is newly within reach, it was not within reach before, at least not at the entry-level price and power tier NVIDIA is targeting with Cognex, Doosan Bobcat, Matic and Wing as its named early partners. My own reading of that gap, not NVIDIA's stated one, is that a meaningful share of the "AI-powered" consumer and light-commercial robots that have shipped over the past several years were making a marketing claim that outran what their compute budget could actually run onboard. Nobody had to prove otherwise, because nothing certifies that claim.

Figure AI is the exception worth naming, and it is worth naming precisely because it chose transparency nobody required of it. Alongside its Index data platform announcement this week, Figure disclosed that its crowdsourced training pipeline logs 373 unique tasks, 1,146 unique objects and 116 unique environments for every 1,000 hours of video collected. That is not a certification. It is a company volunteering the kind of granular evidence a skeptical buyer would need to evaluate an AI claim, in a market where nothing forces anyone to disclose it.

What a real verification standard would actually cost

I want to be honest about why this gap has not closed, because the absence is not simply corporate reluctance.

A neural policy resists the kind of standardized bench test that catches a loose wire or an overheating battery. UL, ISO and IFR built their credibility on repeatable, deterministic tests; a trained model's behavior on held-out data does not reduce to a pass or fail the way a shock hazard does, and none of these bodies has the machine-learning evaluation mandate or staff to build one from scratch. There is also a competitive reason nobody volunteers first. A vendor that opens its model to independent evaluation invites exactly the scrutiny a vendor running a thinner stack has every incentive to avoid, which means voluntary disclosure, Figure's approach this week, tends to come only from companies confident enough in their own numbers to publish them.

That is the tradeoff a buyer has to sit with. Waiting for a formal capability-verification standard means waiting for a problem that is genuinely harder to solve than fire testing. Taking a vendor's "AI-powered" label at face value means buying on a specification that nothing on the market currently checks.

From where I sit running a robotics taxonomy, that gap shows up constantly in how differently companies use the same word. NVIDIA's own early-adopter quotes make the point without meaning to: Wing framed its Jetson Orin Nano 2 evaluation around flight-time efficiency, while Matic Robotics framed the identical chip around adding onboard capability. Two companies, one module, two entirely different claims about what "AI-powered" will mean for their product line by the time it ships.

For a procurement team evaluating any AI-labeled robot over the next year, the useful question has stopped being which safety mark or classification a vendor holds. Ask instead what a vendor discloses about task diversity, held-out performance, or failure rate outside a demo environment, the way Figure did unprompted this week. A UL 3300 seal tells a buyer the robot will not catch fire. It says nothing about what happens behind its eyes.

The next real credential in robotics will not come from a standards body updating a vocabulary list or a regulator requiring a disclosure sentence. It will come from buyers who start asking vendors for the numbers Figure volunteered, and treating a vendor's silence on that question as the answer.

Hero image: Tally autonomous shelf-scanning robot in a US grocery store, official image from Simbe Robotics.

Disclaimer: This article is for general information purposes only and does not constitute investment, legal, or procurement advice. Readers should verify details with primary sources before making business decisions.

robot certificationUL 3300AI verificationrobotics standardsEU AI ActNVIDIAFigure AIprocurement