# Matched‑recipe comparison of pre‑training checkpoints

> All three downstream stages (mid‑training, long‑context adaptation, and SFT) are run with the same learning‑rate recipe for each checkpoint.

Oossa · 2026-09-29 · https://oossa.com/en/matched-recipe-comparison-of-pre-training-checkpoints

## What recipe was chosen

We swept the nine possible ( m, ℓ ) combinations – where m and ℓ are the mid‑training and long‑context learning‑rate factors drawn from {1, 1/3, 1/9} – while keeping the SFT factor fixed at s = 1/3.  The sweep was performed on the **Cooldown** checkpoint, and the combination that gave the highest post‑SFT aggregate benchmark score was **(1, 1)(1, 1)**.

## Why the same recipe is used for all checkpoints

Because the training‑infrastructure constraints prevented us from tuning a separate learning‑rate schedule for each pre‑training checkpoint, we adopted the **(1, 1)(1, 1)** recipe for every checkpoint (Constant, Cooldown, and Merge).  This ensures a *matched* downstream recipe, so any differences in the final scores can be attributed to the underlying pre‑training checkpoint rather than to a different fine‑tuning schedule.

## The facts

- Mid‑training factor m = 1
- Long‑context factor ℓ = 1
- SFT factor s = 1/3
- Chosen because it gave the highest post‑SFT aggregate score on the Cooldown checkpoint.

## Why it matters

Using a single, best‑performing recipe across all checkpoints makes the comparison fair: the only variable is the quality of the pre‑training checkpoint itself, not the downstream learning‑rate schedule.

## Sources & references

1. [Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack](https://www.aleph-alpha.com/en/blog/good-pretraining-bad-sft/) – Aleph Alpha, 2026-09-29

Last updated: 2026-09-29
