Towards Sampling Data Structures for Tensor Products in Turnstile Streams
Song, Zhao, Xie, Shenghao, Zhou, Samson
This paper studies the computational challenges of large-scale attention-based models in artificial intelligence by utilizing importance sampling methods in the streaming setting. Inspired by the classical definition of the $\ell_2$ sampler and the recent progress of the attention scheme in Large Language Models (LLMs), we propose the definition of the attention sampler. Our approach significantly reduces the computational burden of traditional attention mechanisms. We analyze the effectiveness of the attention sampler from a theoretical perspective, including space and update time. Additionally, our framework exhibits scalability and broad applicability across various model architectures and domains.
Oct-7-2025
- Country:
- North America > United States
- Texas (0.04)
- District of Columbia > Washington (0.04)
- California > Alameda County
- Berkeley (0.04)
- Europe
- Italy (0.04)
- Spain
- Canary Islands (0.04)
- Catalonia > Barcelona Province
- Barcelona (0.04)
- Asia
- Singapore (0.04)
- Middle East > Jordan (0.04)
- British Indian Ocean Territory > Diego Garcia (0.04)
- Afghanistan > Parwan Province
- Charikar (0.04)
- North America > United States
- Genre:
- Research Report (1.00)
- Technology: