End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
Goetting, Dylan, Singh, Himanshu Gaurav, Loquercio, Antonio
–arXiv.org Artificial Intelligence
The ability to navigate effectively within an environment to achieve a goal is a hallmark of physical intelligence. Spatial memory, along with more advanced forms of spatial cognition, is believed to have begun evolving early in the history of land animals and advanced vertebrates, likely between 400 and 200 million years ago [1]. Because this ability has evolved over such a long period, it feels almost instinctual and trivial to humans. However, navigation is, in reality, a highly complex problem. It requires the coordination of low-level planning to avoid obstacles alongside high-level reasoning to interpret the environment's semantics and explore the directions that are most likely to get the agent to achieve their goals. A significant portion of the navigation problem appears to involve cognitive processes similar to those required for answering long-context image and video questions, an area where contemporary vision-language models (VLMs) excel [2, 3]. However, when naively applied to navigation tasks, these models face clear limitations. Specifically, when given a task description concatenated with an observation-action history, VLMs often struggle to produce fine-grained spatial outputs to avoid obstacles and fail to effectively utilize their long-context reasoning capabilities to support effective navigation [4, 5, 6].
arXiv.org Artificial Intelligence
Nov-8-2024
- Country:
- North America > United States
- Pennsylvania (0.04)
- California > Alameda County
- Berkeley (0.04)
- Europe > United Kingdom
- England > Cambridgeshire > Cambridge (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.93)
- Technology: