Pretraining Scaling Laws for Generative Evaluations of Language Models

Schaeffer, Rylan, Levi, Noam, Miranda, Brando, Koyejo, Sanmi

arXiv.org Artificial Intelligence 

Neural scaling laws have played a central role in modern machine learning, driving the field's ever-expanding scaling of parameters, data and compute. While much research has gone into fitting scaling laws and predicting performance on pretraining losses and on discriminative evaluations such as multiple-choice question-answering benchmarks, comparatively little research has been done on fitting scaling laws and predicting performance on generative evaluations such as mathematical problem-solving or coding. In this work, we propose and evaluate three different pretraining scaling laws for fitting pass-at-k on generative evaluations and for predicting pass-at-k of the most expensive model using the performance of cheaper models. Our three scaling laws differ in the covariates used: (1) pretraining compute, (2) model parameters and pretraining tokens, (3) log likelihoods of gold reference solutions. We make four main contributions: First, we show how generative evaluations offer new hyperparameters (in our setting, k) that researchers can use to control the scaling laws parameters and the predictability of performance. Second, in terms of scaling law parameters, we find that the compute scaling law and parameters + tokens scaling law stabilize for the last 1.5 2.5 orders of magnitude, whereas the gold reference likelihood scaling law stabilizes for the last 5 orders of magnitude.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found