More Context, Less Distraction: Zero-shot Visual Classification by Inferring and Conditioning on Contextual Attributes

An, Bang, Zhu, Sicheng, Panaitescu-Liess, Michael-Andrei, Mummadi, Chaithanya Kumar, Huang, Furong

Oct-8-2023–arXiv.org Artificial Intelligence

Vision-language models like CLIP are widely used in zero-shot image classification due to their ability to understand various visual concepts and natural language descriptions. However, how to fully leverage CLIP's unprecedented human-like understanding capabilities to achieve better performance is still an open question. This paper draws inspiration from the human visual perception process: when classifying an object, humans first infer contextual attributes (e.g., background and orientation) which help separate the foreground object from the background, and then classify the object based on this information. Inspired by it, we observe that providing CLIP with contextual attributes improves zero-shot image classification and mitigates reliance on spurious features. We also observe that CLIP itself can reasonably infer the attributes from an image. With these observations, we propose a training-free, two-step zero-shot classification method PerceptionCLIP. Given an image, it first infers contextual attributes (e.g., background) and then performs object classification conditioning on them. Our experiments show that PerceptionCLIP achieves better generalization, group robustness, and interpretability. For example, PerceptionCLIP with ViT-L/14 improves the worst group accuracy by 16.5% on the Waterbirds dataset and by 3.5% on CelebA.

classification, contextual, dataset, (16 more...)

arXiv.org Artificial Intelligence

Oct-8-2023

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - Maryland (0.04)
  - New York (0.04)
- Europe > Switzerland
  - Zürich > Zürich (0.14)

Genre:
- Research Report > New Finding (0.46)

Industry:
- Government
  - Military (0.67)
  - Regional Government > North America Government
    - United States Government (0.67)

Technology:
- Information Technology > Artificial Intelligence
  - Vision (1.00)
  - Natural Language > Large Language Model (1.00)
  - Machine Learning > Neural Networks
    - Deep Learning (0.46)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found