NoTVLA: Narrowing of Dense Action Trajectories for Generalizable Robot Manipulation

Huang, Zheng, Liu, Mingyu, Lin, Xiaoyi, Zhu, Muzhi, Zhao, Canyu, Du, Zongze, Li, Xiaoman, Jia, Yiduo, Zhong, Hao, Chen, Hao, Shen, Chunhua

arXiv.org Artificial Intelligence 

Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic forgetting. This issue stems from their overreliance on continuous action sequences or action chunks, which inadvertently create isolated data silos that disrupt knowledge retention across tasks. To tackle these challenges, we propose the Narrowing of Trajectory VLA (NoTVLA) framework: a novel approach that narrows its focus to sparse trajectories, thereby avoiding the catastrophic forgetting associated with dense trajectory fine-tuning. A key innovation of NoTVLA lies in its trajectory planning strategy: instead of centering on the target object's trajectory, it leverages temporal compression and spatial reasoning pruning specifically for the robot end effector's trajectory. Furthermore, training is conducted using these sparse trajectories rather than dense action trajectories, an optimization that delivers remarkable practical advantages with better performance in zero-shot. This design ensures that NoTVLA's operational accuracy closely approximates that of single-task expert models. Crucially, it also preserves the model's inherent language capabilities, enabling zero-shot generalization in specific scenarios, supporting unified model deployment across multiple robot platforms, and fostering a degree of generalization even when perceiving tasks from novel perspectives. Vision-Language-Action (VLA) models have emerged as a transformative force for embodied intelligence, heralding a new era of integrated perception, reasoning, and control (Kim et al., 2024; Brohan et al., 2023; Driess et al., 2023; Alayrac et al., 2022). By coupling large-scale multimodal reasoning engines with action experts, these systems aspire to endow robots with the capacity to operate robustly in open-ended, unstructured environments. Y et, despite rapid progress, several obstacles still impede reliable deployment. Foremost among them is catastrophic forgetting (French, 1999; Kirkpatrick et al., 2017; Zenke et al., 2017), which is aggravated in current pipelines by an overreliance on dense, continuously parameterized action trajectories (or "action chunks"). Such representations foster isolated data silos: when finetuned on a new task, the model may overwrite previously consolidated competencies, degrading prior-task performance.