Caption-Driven Explorations: Aligning Image and Text Embeddings through Human-Inspired Foveated Vision

Zanca, Dario, Zugarini, Andrea, Dietz, Simon, Altstidl, Thomas R., Ndjeuha, Mark A. Turban, Schwinn, Leo, Eskofier, Bjoern

Aug-19-2024–arXiv.org Artificial Intelligence

Understanding human attention is crucial for vision science and AI. While many models exist for free-viewing, less is known about task-driven image exploration. To address this, we introduce CapMIT1003, a dataset with captions and click-contingent image explorations, to study human attention during the captioning task. We also present NevaClip, a zero-shot method for predicting visual scanpaths by combining CLIP models with NeVA algorithms. NevaClip generates fixations to align the representations of foveated visual stimuli and captions. The simulated scanpaths outperform existing human attention models in plausibility for captioning and free-viewing tasks. This research enhances the understanding of human attention and advances scanpath prediction models.

caption, nevaclip, scanpath, (13 more...)

arXiv.org Artificial Intelligence

Aug-19-2024

arXiv.org PDF

Add feedback

Country:
- Europe
  - Germany
    - Bavaria
      - Middle Franconia > Nuremberg (0.05)
      - Upper Bavaria > Munich (0.06)
    - North Rhine-Westphalia > Upper Bavaria
      - Munich (0.05)
  - Italy (0.05)

Genre:
- Research Report (0.51)

Industry:
- Health & Medicine > Therapeutic Area (0.49)

Technology:
- Information Technology > Artificial Intelligence
  - Cognitive Science (1.00)
  - Machine Learning (0.97)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found