The llama.cpp project released version b11474 on 2026‑10‑07. It adds a GLM5‑Next multi‑token‑prediction (MTP) graph, called NextN, to the model runner. The change lets the code skip unnecessary compute when no output rows are needed, dropping the 4‑token catch‑up latency from 6.9 ms to 0.33 ms. The update also fixes extraction contracts and improves how partial recurrent rollbacks are handled.
Why it matters
Developers can now run llama.cpp models faster, especially when using draft‑based generation that relies on multi‑token prediction.