Large language models can do jaw-dropping things. But nobody knows exactly why.

MIT Technology Review 

Grokking is just one of several odd phenomena that have AI researchers scratching their heads. The largest models, and large language models in particular, seem to behave in ways textbook math says they shouldn't. This highlights a remarkable fact about deep learning, the fundamental technology behind today's AI boom: for all its runaway success, nobody knows exactly how--or why--it works. "Obviously, we're not completely ignorant," says Mikhail Belkin, a computer scientist at the University of California, San Diego. "But our theoretical analysis is so far off what these models can do. Like, why can they learn language? I think this is very mysterious."

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found