Balancing Performance and Efficiency in Zero-shot Robotic Navigation
Kuzmenko, Dmytro, Shvai, Nadiya
–arXiv.org Artificial Intelligence
We present an optimization study of the Vision-Language Frontier Maps (VLFM) applied to the Object Goal Navigation task in robotics. Our work evaluates the efficiency and performance of various vision-language models, object detectors, segmentation models, and multi-modal comprehension and Visual Question Answering modules. Using the $\textit{val-mini}$ and $\textit{val}$ splits of Habitat-Matterport 3D dataset, we conduct experiments on a desktop with limited VRAM. We propose a solution that achieves a higher success rate (+1.55%) improving over the VLFM BLIP-2 baseline without substantial success-weighted path length loss while requiring $\textbf{2.3 times}$ less video memory. Our findings provide insights into balancing model performance and computational efficiency, suggesting effective deployment strategies for resource-limited environments.
arXiv.org Artificial Intelligence
Jun-5-2024
- Country:
- Europe > Ukraine
- Kyiv Oblast > Kyiv (0.05)
- Asia > South Korea
- Europe > Ukraine
- Genre:
- Research Report > New Finding (0.89)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Robots (1.00)
- Natural Language > Large Language Model (0.68)
- Information Technology > Artificial Intelligence