Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization

Open in new window