The Download: the next big thing in LLMs and how AI academic research is shifting

MIT Technology Review 

Plus: Nvidia has secured $500 billion from Wall Street for AI infrastructure. Nine years after Google researchers introduced the transformer, this family of neural networks has become the engine inside every major large language model. But transformers are starting to show their age. As LLMs get bigger and better, transformers have become a bottleneck. Their dense attention mechanism becomes increasingly expensive as the amount of text grows, and they're not great at keeping track of a lot of information at once. Here are four new ideas for how to solve the transformer problem --innovations that could change LLMs for good, making them faster, far more efficient, and (maybe) even smarter.