Improving Pretraining Data Using Perplexity Correlations
Thrush, Tristan, Potts, Christopher, Hashimoto, Tatsunori
Quality pretraining data is often seen as the key to high-performance language models. However, progress in understanding pretraining data has been slow due to the costly pretraining runs required for data selection experiments. We present a framework that avoids these costs and selects high-quality pretraining data without any LLM training of our own. Our work is based on a simple observation: LLM losses on many pretraining texts are correlated with downstream benchmark performance, and selecting high-correlation documents is an effective pretraining data selection method. We build a new statistical framework for data selection centered around estimates of perplexity-benchmark correlations and perform data selection using a sample of 90 LLMs taken from the Open LLM Leaderboard on texts from tens of thousands of web domains. In controlled pretraining experiments at the 160M parameter scale on 8 benchmarks, our approach outperforms DSIR on every benchmark, while matching the best data selector found in DataComp-LM, a hand-engineered bigram classifier. Dataset curation is increasingly crucial for training high-quality large language models (LLMs). However, data-driven approaches to data selection typically involve expensive model retraining steps that limit their effectiveness, and no algorithm has been reported to consistently beat or match hand-crafted classifiers for data selection (Li et al., 2024). Is training new LLMs necessary for data selection? Instead of training our own models, can we use the growing collection of publicly available, high-performance LLMs (Wolf et al., 2019; Beeching et al., 2023) to perform data valuation and selection? This would have significant benefits: we could leverage the millions of dollars collectively spent on building these LLMs, and we would have coverage over a large, heterogeneous collection of high-performance models varying in size, architectures, and pretraining data distribution. Despite these advantages, using existing models for pretraining data selection is challenging, as the training data for these models are often unknown and heterogeneous.
Sep-9-2024
- Country:
- North America > United States
- Europe > Italy
- Calabria > Catanzaro Province > Catanzaro (0.04)
- Asia > Middle East
- Jordan (0.04)
- Genre:
- Research Report
- Experimental Study (0.66)
- New Finding (0.46)
- Research Report
- Technology: