OossaAI is evolving fast. We explain it simply.
Newsletter

Note · 1 min read

Matched‑recipe comparison of pre‑training checkpoints

All three downstream stages (mid‑training, long‑context adaptation, and SFT) are run with the same learning‑rate recipe for each checkpoint.

Oossa · About Oossa

What recipe was chosen

We swept the nine possible ( m, ℓ ) combinations – where m and ℓ are the mid‑training and long‑context learning‑rate factors drawn from {1, 1/3, 1/9} – while keeping the SFT factor fixed at s = 1/3. The sweep was performed on the **Cooldown** checkpoint, and the combination that gave the highest post‑SFT aggregate benchmark score was **(1, 1)(1, 1)**.

Why the same recipe is used for all checkpoints

Because the training‑infrastructure constraints prevented us from tuning a separate learning‑rate schedule for each pre‑training checkpoint, we adopted the **(1, 1)(1, 1)** recipe for every checkpoint (Constant, Cooldown, and Merge). This ensures a *matched* downstream recipe, so any differences in the final scores can be attributed to the underlying pre‑training checkpoint rather than to a different fine‑tuning schedule.

Why it matters

Using a single, best‑performing recipe across all checkpoints makes the comparison fair: the only variable is the quality of the pre‑training checkpoint itself, not the downstream learning‑rate schedule.

Was this article useful?

Sources & references

#SourceOutletDateKey takeaway
1Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack ↗Aleph AlphaSep 29, 2026

1 sources

Last updated:

Oossallms.txt.md

Share

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.