Improving Visual Reasoning by Exploiting The Knowledge in Texts
Sharifzadeh, Sahand, Baharlou, Sina Moayed, Schmitt, Martin, Schütze, Hinrich, Tresp, Volker
–arXiv.org Artificial Intelligence
This paper presents a new framework for training image-based classifiers from a combination of texts and images with very few labels. We consider a classification framework with three modules: a backbone, a relational reasoning component, and a classification component. While the backbone can be trained from unlabeled images by self-supervised learning, we can fine-tune the relational reasoning and the classification components from external sources of knowledge instead of annotated images. By proposing a transformer-based model that creates structured knowledge from textual input, we enable the utilization of the knowledge in texts. We show that, compared to the supervised baselines with 1% of the annotated images, we can achieve ~8x more accurate results in scene graph classification, ~3x in object classification, and ~1.5x in predicate classification.
arXiv.org Artificial Intelligence
Feb-9-2021
- Country:
- North America > United States
- California (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Hawaii > Honolulu County
- Honolulu (0.04)
- Europe
- Spain > Valencian Community
- Valencia Province > Valencia (0.04)
- Italy > Lazio
- Rome (0.04)
- Germany
- Berlin (0.04)
- Bavaria > Upper Bavaria
- Munich (0.04)
- Spain > Valencian Community
- Asia > Middle East
- Republic of Türkiye > Karaman Province > Karaman (0.04)
- North America > United States
- Genre:
- Research Report (0.82)
- Technology: