A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
Han, ByungOk, Kim, Jaehong, Jang, Jinhyeok
–arXiv.org Artificial Intelligence
Vision-Language-Action (VLA) models are designed to enable robots to generate actions based on a user's task instruction by following three key steps: (1) interpreting the task instruction, (2) analyzing the current visual information in relation to the task, and (3) predicting the necessary actions for execution. By combining vision and language inputs, VLA models allow robots to perform complex tasks using both visual context and linguistic commands. Recently, Large Language Models (LLMs) [1, 2, 3] and Vision-Language Models (VLMs) [4, 5, 6] have reported high capabilities to general understanding. VLA models have leveraged VLMs to enhance a robot's perception capabilities, showing promising results in their ability to interpret and execute complex tasks. By this, recent VLA have demonstrated accurate action generation across various tasks, utilizing diverse robot hardware in real-world environments such as RT-2 [7], RoboFlamingo [8], OpenVLA [9], LLaRA [10], and LLARVA [11].
arXiv.org Artificial Intelligence
Oct-20-2024
- Country:
- Europe
- Netherlands > South Holland
- Delft (0.04)
- Germany > Bavaria
- Upper Bavaria > Munich (0.04)
- Netherlands > South Holland
- Europe
- Genre:
- Research Report (0.50)
- Technology: