Oossa

llama.cpp adds GLM5‑Next multi‑token prediction and speed tweaks

The open‑source LLM runner now supports a new GLM5‑Next MTP head, cutting 4‑token catch‑up time to 0.33 ms.

NoteBy Published by Oossa: 1 min read

The llama.cpp project released version b11474 on 2026‑10‑07. It adds a GLM5‑Next multi‑token‑prediction (MTP) graph, called NextN, to the model runner. The change lets the code skip unnecessary compute when no output rows are needed, dropping the 4‑token catch‑up latency from 6.9 ms to 0.33 ms. The update also fixes extraction contracts and improves how partial recurrent rollbacks are handled.

Why it matters

Developers can now run llama.cpp models faster, especially when using draft‑based generation that relies on multi‑token prediction.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.