Goto

Collaborating Authors

 Deep Learning


A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP) Weijie T u 1 Weijian Deng

Neural Information Processing Systems

Contrastive Language-Image Pre-training (CLIP) models have demonstrated remarkable generalization capabilities across multiple challenging distribution shifts. However, there is still much to be explored in terms of their robustness to the variations of specific visual factors.



99cad265a1768cc2dd013f0e740300ae-Paper.pdf

Neural Information Processing Systems

Graph Neural Networks (GNNs) have been an active research field for the last ten years with significant advancements in graph representation learning [1, 2, 3, 4].


TransGAN: TwoPureTransformersCanMakeOne StrongGAN,andThatCanScaleUp

Neural Information Processing Systems

The recent explosive interest on transformers has suggested their potential to become powerful "universal" models for computer vision tasks, such as classification, detection, and segmentation. While those attempts mainly study the discriminativemodels, weexplore transformers onsome more notoriously difficult vision tasks, e.g., generative adversarial networks (GANs). Our goal is to conduct the first pilot study in building a GANcompletely free of convolutions, using only pure transformer-based architectures.


Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint Selection

Neural Information Processing Systems

Given the complexities inherent in visual scenes, such as object occlusion, a comprehensive understanding often requires observation from multiple viewpoints. Existing multi-viewpoint object-centric learning methods typically employ random or sequential viewpoint selection strategies.