NaVIP: An Image-Centric Indoor Navigation Solution for Visually Impaired People

Yu, Jun, Zhang, Yifan, Aila, Badrinadh, Namboodiri, Vinod

Oct-8-2024–arXiv.org Artificial Intelligence

Indoor navigation is challenging due to the absence of satellite positioning. This challenge is manifold greater for Visually Impaired People (VIPs) who lack the ability to get information from wayfinding signage. Other sensor signals (e.g., Bluetooth and LiDAR) can be used to create turn-by-turn navigation solutions with position updates for users. Unfortunately, these solutions require tags to be installed all around the environment or the use of fairly expensive hardware. Moreover, these solutions require a high degree of manual involvement that raises costs, thus hampering scalability. We propose an image dataset and associated image-centric solution called NaVIP towards visual intelligence that is infrastructure-free and task-scalable, and can assist VIPs in understanding their surroundings. Specifically, we start by curating large-scale phone camera data in a four-floor research building, with 300K images, to lay the foundation for creating an image-centric indoor navigation and exploration solution for inclusiveness. Every image is labelled with precise 6DoF camera poses, details of indoor PoIs, and descriptive captions to assist VIPs.

large language model, machine learning, natural language, (16 more...)

arXiv.org Artificial Intelligence

Oct-8-2024

arXiv.org PDF

Add feedback

Country:
- North America > Puerto Rico
  - Peñuelas > Peñuelas (0.04)
- Europe
  - Greece (0.04)
  - Netherlands > North Holland
    - Amsterdam (0.04)

Genre:
- Research Report (1.00)

Industry:
- Health & Medicine (1.00)

Technology:
- Information Technology
  - Sensing and Signal Processing > Image Processing (1.00)
  - Communications > Mobile (1.00)
  - Human Computer Interaction (0.93)
  - Data Science (0.92)
  - Artificial Intelligence
    - Vision (1.00)
    - Robots (1.00)
    - Representation & Reasoning (1.00)
    - Natural Language > Large Language Model (0.94)
    - Machine Learning > Neural Networks
      - Deep Learning (1.00)