Research on Navigation Methods Based on LLMs
–arXiv.org Artificial Intelligence
The emergence of large language models (LLMs) has revolutionized autonomous planning through their exceptional capacity to comprehend and reason about complex, context-rich scenarios [18]. Recent advancements such as NA VGPT and NA VGPT2 convert visual scene semantics into LLM-compatible input prompts, enabling visual-language navigation through LLMs' commonsense knowledge and reasoning capabilities. However, current LLM-based implementations like VLTNet with Tree-of-Thought networks for language-driven zero-shot object navigation (L-ZSON) [14], UniGoal [19] with unified graph representations for general zero-shot navigation, and MapNav's end-to-end visual-language navigation using annotated semantic maps (ASM) as historical frame replacements, while leveraging LLM capabilities, fail to fully exploit robotic systems' inherent potential. These approaches exhibit limited adaptability to diverse environmental conditions and do not enhance robots' intrinsic generalization capabilities. Our proposed methodology addresses these limitations through two key innovations: First, we decompose robotic functionalities into modu-Corresponding author: jianmin@ustc.edu.cn
arXiv.org Artificial Intelligence
Apr-23-2025