Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
Howell, Owen Lewis, Zhao, Linfeng, Zhu, Xupeng, Qian, Yaoyao, Huang, Haojie, Sun, Lingfeng, Thomason, Wil, Platt, Robert, Walters, Robin
–arXiv.org Artificial Intelligence
Transformer-based models have demonstrated effectiveness beyond language processing, showing strong performance in geometry-aware tasks such as robotics, structural biochemistry, and materials science [1-5]. For instance, 3D robotic perception tasks ranging from segmentation to object matching process point clouds and LiDAR data using attention mechanisms. These tasks heavily rely on token-based representations, and their performance is often constrained by the number of tokens the model can effectively handle. AlphaFold [6], for example, employs equivariant transformers to predict protein structures with unprecedented accuracy by explicitly leveraging SE (3) symmetries such as rotations and translations. However, implementing an equivariant neural network structure typically incurs significant computational overhead and increased inference time. As a result, most current approaches are limited to small symmetry groups or low-order representations [7-12]. Enabling fast, low-memory overhead equivariant operations over large context windows is essential to scaling robust and sample-efficient learning in geometry-aware domains. Unfortunately, maintaining equivariance while modeling a global geometric context is challenging due to the computational demands of processing high dimensional data at scale. There are essentially two components that contribute to the computational complexity of E(3)-equivariant transformers: the time and memory scaling of the transformer with the number of tokens, N, and the time and memory complexity on the maximum harmonic degree, ℓ.
arXiv.org Artificial Intelligence
Sep-30-2025
- Country:
- North America > United States (0.28)
- Genre:
- Research Report (0.52)
- Industry:
- Information Technology (0.66)
- Technology: