Grounding Everything: Emerging Localization Properties in Vision-Language Transformers

Open in new window