Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
Mane, Atharv Mahesh, Weerakoon, Dulanga, Subbaraju, Vigneshwaran, Sen, Sougata, Sarma, Sanjay E., Misra, Archan
–arXiv.org Artificial Intelligence
Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. T o address this gap, we introduce a data augmentation framework-Imputer, and use it to curate a new benchmark dataset-ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. W e also propose Ges3ViG, a novel model for 3D-ERU that achieves 30% improvement in accuracy as compared to other 3D-ERU models and 9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/
arXiv.org Artificial Intelligence
Apr-15-2025
- Country:
- North America > United States (0.46)
- Genre:
- Research Report > Promising Solution (0.48)
- Technology: