A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM

Han, ByungOk, Kim, Jaehong, Jang, Jinhyeok

arXiv.org Artificial Intelligence 

Vision-Language-Action (VLA) models are designed to enable robots to generate actions based on a user's task instruction by following three key steps: (1) interpreting the task instruction, (2) analyzing the current visual information in relation to the task, and (3) predicting the necessary actions for execution. By combining vision and language inputs, VLA models allow robots to perform complex tasks using both visual context and linguistic commands. Recently, Large Language Models (LLMs) [1, 2, 3] and Vision-Language Models (VLMs) [4, 5, 6] have reported high capabilities to general understanding. VLA models have leveraged VLMs to enhance a robot's perception capabilities, showing promising results in their ability to interpret and execute complex tasks. By this, recent VLA have demonstrated accurate action generation across various tasks, utilizing diverse robot hardware in real-world environments such as RT-2 [7], RoboFlamingo [8], OpenVLA [9], LLaRA [10], and LLARVA [11].

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found