Back to Search
W

WITA-Omni

ModelActive

WITA-Omni is AGIBOT's embodied-native omni-modal model designed for real-world human-robot interaction, publicly demonstrated on 2026-09-21 at The Greater Bay Area Film Concert, where an AGIBOT robot held live face-to-face interaction with a celebrity, following conversations, identifying who to respond to, and coordinating speech, movement, and expression in real time without scripted sequences. Unlike multimodal AI systems built to process combinations of text, images, and audio, WITA-Omni is built for continuous, unpredictable physical-world interaction: it brings visual, audio, language, and temporal information into a unified framework, allowing the model to follow an interaction over time and understand who said or did what, when it happened, and how different signals relate. The model is trained on large-scale multimodal data centered on real-world human interaction, with training data preserving the timing and relationships between speech, visual information, language, movement, and expression, refined by a multi-stage training process targeting interaction decision-making. At its core is a Thinker-Talker-Actor architecture, extending the Thinker-Talker framework with an Actor designed specifically for embodied interaction: the Thinker interprets what is happening and makes high-level interaction decisions, the Talker generates speech, and the Actor controls movement and expression. Speech and movement are aligned along the same timeline, so a robot can speak, move, and express itself as part of a single coordinated response, with gestures and expressions adjusting to speech rhythm, pauses, emphasis, and emotional cues. On benchmarks, WITA-Omni Preview ranked first overall on the DailyOmni benchmark with an average accuracy of 85.21% and first place in six of eight evaluation dimensions (company-reported, not independently verified). WITA-Omni is developed by AGIBOT, alongside its line of humanoid robots and embodied AI models such as GE-Act 2.0.

Details

Updated:9/25/2026

No structured details available.

Tags

thinker-talker-actorspeechhuman-robot-interactionagibotembodied-aiomni-modal

Relationships

Sources

AGIBOT Demonstrates WITA-Omni in Live Human-Robot Interaction at The Greater Bay Area Film Concert
website
Visit

Appears In

No related landscapes.