Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation
Somerstep, Seamus, Raman, Vinod, Subedi, Unique, Sun, Yuekai
Using the bit string generation problem as a case study, we th eoretically compare two standard methods for adapting large language models to n ew tasks. The first, referred to as supervised fine-tuning, involves training a new next token predictor on good generations. The second method, Best-of-N, trains a reward model to select good responses from a collection generated by an unal tered base model. If the learning setting is realizable, we find that supervised fi ne-tuning outperforms BoN through a better dependence on the response length in its rate of convergence. If realizability fails, then depending on the failure mode, BoN can enjoy a better rate of convergence in either n or a rate of convergence with better dependence on the response length.
May-26-2025
- Country:
- North America > United States
- Michigan (0.04)
- Europe > United Kingdom
- England
- Oxfordshire > Oxford (0.04)
- Cambridgeshire > Cambridge (0.04)
- England
- Asia > Middle East
- Jordan (0.04)
- North America > United States
- Genre:
- Research Report > Experimental Study (0.67)
- Technology: