Progressive Residual Warmup for Language Model Pretraining Paper • 2603.05369 • Published 7 days ago • 31 • 5