Lightweight Vision Transformer with Bidirectional Interaction

Neural Information Processing Systems 

Recent advancements in vision backbones have significantly improved their performance by simultaneously modeling images' local and global contexts. However, the bidirectional interaction between these two contexts has not been well explored and exploited, which is important in the human visual system.