TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

Zhang, Zongzheng, Xu, Haobo, Yang, Zhuo, Yue, Chenghao, Lin, Zehao, Gao, Huan-ang, Wang, Ziwei, Zhao, Hao

arXiv.org Artificial Intelligence 

Understanding physical interactions through force cues is essential for mastering real-world robotic manipulation. One particularly informative signal is joint torque, which reflects subtle variations in end-effector contact dynamics without requiring external force sensors [1, 2, 3]. As shown in Figure 1(a), different outcomes in a seemingly simple task like charger insertion--no contact, failed insertion, and successful plug-in--can be clearly distinguished by the joint torque profiles of a 7-DoF arm. These torque responses offer rich physical context that is otherwise imperceptible from RGB observations alone. However, despite the growing success of Vision-Language-Action (VLA) models [4, 5, 6, 7, 8] in bridging vision and control, their ability to interpret and leverage such physical feedback remains limited. Our work aims to bridge this gap by integrating torque signals into pretrained VLA models, enabling contact-sensitive decision-making without compromising generalization or scalability. The challenge lies in how to embed torque into VLA architectures. Torque is a proprioceptive signal, structurally different from image and language inputs, and varies across time, especially during contact-rich phases. As illustrated in Figure 1(c), multiple torque integration strategies exist across three axes--when (immediate vs. historical vs. predictive), where (encoder vs. decoder), and