Goto

Collaborating Authors

 Large Language Model









Exposing Attention Glitches with Flip-Flop Language Modeling

Neural Information Processing Systems

This simple generative task requires a model to copy binary symbols over long-range dependencies, ignoring the tokens in between. We find that Transformer FFLMs suffer from a long tail of sporadic reasoning errors, some of which we can eliminate using various regularization techniques.



Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIP

Neural Information Processing Systems

Vision-language pre-training methods, e.g., CLIP, demonstrate an impressive zero-shot performance on visual categorizations with the class proxy from the text embedding of the class name.