RobotAIGeek

AGIBOT's WITA-Omni Holds an Unscripted Conversation On Stage

AGIBOT demonstrated WITA-Omni, an embodied AI model built for real-time interaction timing, letting a humanoid hold an unscripted live conversation with a TV host in Macau, a test of whether interactive robots are ready for service settings.

martti
2 Min. LesezeitPosted: 22. Sept. 2026
AGIBOT's WITA-Omni Holds an Unscripted Conversation On Stage

At a televised concert in Macau on Monday, September 21, a humanoid robot built by Shanghai-based AGIBOT sat across from a Chinese entertainment host and simply talked with her, following the flow of a live, unrehearsed exchange rather than executing a fixed sequence of pre-programmed lines. The robot tracked who was speaking, waited for a natural opening to respond, and matched its speech, gestures and posture to the rhythm of the conversation in real time. The demonstration, staged at the Greater Bay Area Film Concert, was AGIBOT's first large-audience showcase of WITA-Omni, an embodied-native AI model built specifically to let robots follow and participate in unscripted human interaction rather than react to isolated commands.

AGIBOT, formally AGIBOT Innovation (Shanghai) Technology Co., Ltd. and also known in China as Zhiyuan Robotics, was founded in February 2023 by former Huawei engineers Deng Taihua and Peng Zhihui and has since become one of the more heavily funded entrants in China's humanoid-robot field, reporting that its 15,000th robot rolled off the production line in June 2026 and opening a Hong Kong IPO process that July. Its product line spans the A-series and X-series humanoids, the G-series general-purpose robot, D-series quadrupeds and a commercial cleaning platform, giving WITA-Omni a wide set of hardware to run on if AGIBOT chooses to roll the model out beyond a single stage demonstration.

Turning Multimodal Perception Into a Live Exchange

Most of the multimodal AI systems now shipping in robotics and consumer devices are built to process combinations of text, image and audio input and return an answer. That is a meaningfully different problem from holding a live conversation in a room full of people, where a system also has to work out who is speaking, what was said a few seconds earlier, whom a gesture or comment was directed at, and whether the moment calls for a response at all. WITA-Omni is built around that second problem. The model fuses visual, audio, language and timing signals into a single framework so it can track an interaction as it unfolds rather than treating each input as a discrete query, and AGIBOT has published benchmark results to back the claim: on the third-party DailyOmni audio-visual reasoning benchmark, a preview version of WITA-Omni posted an average accuracy of 85.21 percent, finishing first overall and topping six of eight evaluation categories against general-purpose multimodal systems from other AI labs.

The architecture behind that score is what AGIBOT calls Thinker-Talker-Actor, an extension of the Thinker-Talker structure used in some conversational AI systems with a third component built specifically for embodiment. The Thinker module interprets what is happening in the scene and decides how the robot should respond at a high level; the Talker generates the actual speech; the Actor governs the robot's physical movement and facial or body expression. Rather than running those three functions as separate, sequential steps, the pipeline aligns them along the same timeline, so a robot's gesture, expression and voice shift together as it speaks instead of arriving as three disconnected outputs stitched together after the fact. AGIBOT trained the model on large-scale multimodal data captured from real human interaction, preserving the timing relationships between speech, movement and expression so the system learns not just what to say but when to say it and how to carry the moment physically.

Why an Unscripted Demo Is a Different Claim Than a Choreographed One

Humanoid robot demonstrations have become routine at trade shows and product launches across China's crowded embodied-AI sector this year, and the overwhelming majority of them are choreographed: a robot walks a fixed path, performs a set gesture, delivers a pre-written line on cue. That format proves a robot can execute a sequence reliably, which matters for manufacturing and logistics work, but it says little about how a robot behaves when a room full of people are not following a script. The Macau demonstration was built to test the opposite condition. The celebrity guest was not reciting lines written for the robot to respond to; the robot had to identify when she was addressing it, judge when a pause in conversation was an invitation to speak rather than an interruption, and adjust its posture and expression to match a tone it had no way of pre-loading. For a procurement team evaluating humanoid robots for reception desks, retail floors, hotel lobbies or entertainment venues, that is the specific capability gap that separates a robot that can perform from a robot that can actually staff a role involving live public interaction.

That distinction is becoming commercially relevant faster than the underlying technology has matured. AGIBOT's own framing of WITA-Omni's applications runs directly at service settings, entertainment venues, retail spaces, hotels and reception areas, where interactions are inherently difficult to script and a robot that freezes or responds off-beat is a worse outcome than a robot that simply performs a rehearsed routine. Competing labs are chasing the same capability from different angles: Figure AI has paired its humanoids with OpenAI's conversational models for similar real-time dialogue work, and Unitree and UBTECH have both pushed voice-interaction features into their consumer and service-oriented product lines this year. What distinguishes AGIBOT's approach is the explicit architectural commitment to timing and turn-taking as first-class problems rather than features layered onto an existing multimodal model after the fact.

The Gap Between a Benchmark Score and a Reliable Deployment

An 85.21 percent accuracy score on a third-party benchmark, even a first-place finish, is not the same claim as a robot ready to work an unsupervised shift in a hotel lobby. Benchmark evaluations like DailyOmni test a model's ability to reason across audio, video and text inputs under controlled conditions; a live, unscripted stage interaction in front of an audience is closer to the deployment reality a buyer actually cares about, but it is still a single, curated demonstration rather than a sustained operational record. AGIBOT has not disclosed failure rates, latency figures under real-world noise conditions, or how WITA-Omni performs when a conversation involves more than one person speaking over each other, a common condition in the retail and hospitality settings the company is targeting. A buyer weighing WITA-Omni-equipped robots against scripted alternatives should treat Monday's demonstration as evidence that the underlying architecture works in at least one live setting, not as proof that the capability is ready for unsupervised commercial deployment at scale.

What the Underlying Data Strategy Signals About AGIBOT's Roadmap

WITA-Omni's training approach, building a model on data that preserves the exact timing relationships between speech, movement and expression, requires a different data-collection pipeline than the video and image datasets most embodied-AI labs rely on for manipulation tasks. That is a meaningful investment decision: AGIBOT is betting that interaction quality, not just task-completion accuracy, will become a differentiator as humanoid robots move out of factories and into settings where they work alongside the public rather than alone on a production line. It also reflects the company's broader manufacturing position. A company that has already shipped 15,000 robots and is pursuing a Hong Kong listing has more incentive than an early-stage startup to build product features that expand the addressable market for hardware it can already produce at volume, rather than to demonstrate a capability with no near-term path to a shipping product.

Reading the Interaction Race Against China's Wider Humanoid Field

China's humanoid-robot manufacturers have spent most of 2026 competing on production volume and mobility benchmarks, running speed records, obstacle courses and factory-floor endurance trials to prove their hardware can survive sustained industrial use. WITA-Omni represents a different competitive axis, one where the differentiator is not how fast or how long a robot can work but how convincingly it can occupy a room with people who are not following a script. That shift matters commercially because the deployment settings AGIBOT is targeting, hospitality, retail and entertainment, pay for a different kind of reliability than a warehouse floor does. A quadruped inspection robot that occasionally misreads a sensor reading can be recalibrated overnight; a reception robot that responds to the wrong guest, or stalls mid-conversation in front of a paying customer, produces a visible and immediate failure that undermines the case for using a robot in that role at all. AGIBOT's decision to build and publicly demonstrate an interaction-timing architecture, rather than simply adding a voice assistant on top of an existing humanoid platform, suggests the company sees that reliability gap as the next competitive front in a market where basic locomotion and manipulation benchmarks are becoming table stakes rather than differentiators.

For a buyer comparing AGIBOT against Chinese rivals such as Unitree and UBTECH, both of which have layered voice-interaction features onto existing product lines this year, the relevant question is not which company has a chatbot running on a humanoid chassis but which one has built turn-taking, interruption-handling and multi-party awareness into the underlying model architecture from the outset. AGIBOT's Thinker-Talker-Actor design is a bet that this distinction will show up in deployment reliability once robots move from single-guest demonstrations into environments with multiple simultaneous speakers, background noise and the ordinary chaos of a working retail floor or hotel lobby, conditions Monday's tightly produced concert setting did not have to contend with.

AGIBOT said it will continue developing WITA-Omni and extending the model across its robotic platforms, with the stated goal of moving robots beyond simply receiving instructions toward taking part in ongoing human interactions as they happen. Whether that ambition survives contact with a real hotel lobby, a crowded retail floor during a holiday rush, or a reception desk fielding three simultaneous questions is a test no single stage demonstration can settle. What Monday's performance in Macau did settle is narrower but still notable: a humanoid robot held its own, unscripted, opposite a professional performer, in front of a live broadcast audience, and did not need a script to do it.

This account synthesizes company disclosures and public reporting on AGIBOT's September 2026 product demonstration. It is for general information purposes only and does not constitute investment, financial, or legal advice.

Hero image credit: AGIBOT.