Note · 1 min read
Matched‑recipe comparison of pre‑training checkpoints
All three downstream stages (mid‑training, long‑context adaptation, and SFT) are run with the same learning‑rate recipe for each checkpoint.
Oossa · About Oossa
What recipe was chosen
We swept the nine possible ( m, ℓ ) combinations – where m and ℓ are the mid‑training and long‑context learning‑rate factors drawn from {1, 1/3, 1/9} – while keeping the SFT factor fixed at s = 1/3. The sweep was performed on the **Cooldown** checkpoint, and the combination that gave the highest post‑SFT aggregate benchmark score was **(1, 1)(1, 1)**.
Why the same recipe is used for all checkpoints
Because the training‑infrastructure constraints prevented us from tuning a separate learning‑rate schedule for each pre‑training checkpoint, we adopted the **(1, 1)(1, 1)** recipe for every checkpoint (Constant, Cooldown, and Merge). This ensures a *matched* downstream recipe, so any differences in the final scores can be attributed to the underlying pre‑training checkpoint rather than to a different fine‑tuning schedule.
Why it matters
Using a single, best‑performing recipe across all checkpoints makes the comparison fair: the only variable is the quality of the pre‑training checkpoint itself, not the downstream learning‑rate schedule.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack ↗ | Aleph Alpha | Sep 29, 2026 |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.