RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
Yao, Yiqun, Li, Xiang, Jiang, Xin, Fang, Xuezhi, Yu, Naitong, Sun, Aixin, Wang, Yequan
–arXiv.org Artificial Intelligence
Humans naturally process real-world multimodal information in a full-duplex manner. In artificial intelligence, replicating this capability is essential for advancing model development and deployment, particularly in embodied contexts. The development of multimodal models faces two primary challenges: (1) effectively handling more than three modalities-such as vision, audio, and text; and (2) delivering full-duplex responses to rapidly evolving human instructions. To facilitate research on models that support both omnimodal processing and full duplexity, we present RoboEgo (alias: FLM-Ego), a unified model system designed to address both challenges. RoboEgo incorporates a backbone architecture and algorithms that natively support full duplexity, achieving a theoretical duplex latency of 80 ms. In streaming visually grounded conversations under real-world conditions, RoboEgo exhibits superior responsiveness and speech naturalness, while maintaining comparable content qualities to state-of-the-art semi-duplex omnimodal models-a feat previously considered unattainable by native full-duplex systems.
arXiv.org Artificial Intelligence
Jun-3-2025
- Genre:
- Research Report (0.83)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Speech (1.00)
- Natural Language
- Large Language Model (1.00)
- Chatbot (1.00)
- Machine Learning > Neural Networks
- Deep Learning (1.00)
- Information Technology > Artificial Intelligence