SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models

Neural Information Processing Systems 

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT) to enhance VLMs' spatial perception and reasoning capabilities.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found