Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
Kobayashi, Masato, Buamanee, Thanpimon
–arXiv.org Artificial Intelligence
Abstract-- We propose Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation (Bi-VLA), a novel framework that extends bilateral control-based imitation learning to handle more than one task within a single model. Conventional bilateral control methods exploit joint angle, velocity, torque, and vision for precise manipulation but require task-specific models, limiting their generality. Bi-VLA overcomes this limitation by utilizing robot joint angle, velocity, and torque data from leader-follower bilateral control with visual features and natural language instructions through SigLIP and FiLM-based fusion. Real-robot experiments showed that Bi-VLA successfully interprets vision-language combinations and improves task success rates compared to conventional bilateral control-based imitation learning. Our Bi-VLA addresses the single-task limitation of prior bilateral approaches and provides empirical evidence that combining vision and language significantly enhances versatility. I. INTRODUCTION Robotic manipulation is increasingly important in human-centered applications such as cooking, eldercare, and interactive service robots [1], [2], [3], [4].
arXiv.org Artificial Intelligence
Sep-24-2025