These startups are chasing the next big thing in LLMs

MIT Technology Review 

Way back in the summer of 2017, AI researchers at Google put out a paper called "Attention Is All You Need," in which they described a new type of neural network called a transformer. It proved to be very good at processing long sequences of data, especially text. Nine years on, transformers are the engines inside every major large language model on the market. "The entire AI industry is built on transformers," says Justin Dangel, cofounder and CEO of the AI startup Subquadratic. "They are one of the most important innovations in the history of computer science, and they've changed the world." But transformers are starting to show their age. Many of the recent advances in LLMs, such as the development of so-called reasoning models and their ability to handle large amounts of input at once, are not neat extensions of that core technology but workarounds that patch over some of its fundamental flaws. A growing number of scientists and engineers are now asking what's coming next. LLMs are not going anywhere, but the way they get built is up for grabs.