The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
Edelman, Benjamin L., Edelman, Ezra, Goel, Surbhi, Malach, Eran, Tsilivis, Nikolaos
–arXiv.org Artificial Intelligence
Large language models (LLMs) exhibit a remarkable ability to perform in-context learning (ICL): learning from patterns in their input context (Brown et al., 2020; Dong et al., 2022). The ability of LLMs to adaptively learn from context is profoundly useful, yet the underlying mechanisms of this emergent capability are not fully understood. In an effort to better understand ICL, some recent works propose to study ICL in controlled synthetic settings--in particular, training transformers on mathematically defined tasks which require learning from the input context. For example, a recent line of works studies the ability of transformers to perform ICL of standard supervised learning problems such as linear regression (Garg et al., 2022; Akyürek et al., 2022; Li et al., 2023; Wu et al., 2023). Studying these well-understood synthetic learning tasks enables fine-grained control over the data distribution, allows for comparisons with established supervised learning algorithms, and facilitates the examination of the in-context "algorithm" implemented by the network. That said, these supervised settings are reflective specifically of few-shot learning, which is only a special case of the more general phenomenon of networks incorporating patterns from their context into their predictions. A few recent works (Bietti et al., 2023; Xie et al., 2022) go beyond the case of cleanly separated in-context inputs and outputs, studying in-context learning on distributions based on discrete stochastic processes. Work done while visiting Harvard University.
arXiv.org Artificial Intelligence
Feb-20-2024
- Country:
- North America > United States (0.93)
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Education (0.34)
- Technology: