# llama.cpp adds GLM5‑Next multi‑token prediction and speed tweaks

> The open‑source LLM runner now supports a new GLM5‑Next MTP head, cutting 4‑token catch‑up time to 0.33 ms.

Oossa · 2026-10-07 · https://oossa.com/en/llama-cpp-adds-glm5-next-multi-token-prediction-and-speed-tweaks

The llama.cpp project released version b11474 on 2026‑10‑07. It adds a GLM5‑Next multi‑token‑prediction (MTP) graph, called NextN, to the model runner. The change lets the code skip unnecessary compute when no output rows are needed, dropping the 4‑token catch‑up latency from 6.9 ms to 0.33 ms. The update also fixes extraction contracts and improves how partial recurrent rollbacks are handled.

## The facts

- Version b11474 released on 2026‑10‑07 ("Wed Oct 07 2026 15:42:20 GMT+0200")
- 4‑token catch‑up latency reduced from 6.9 ms to 0.33 ms

## Why it matters

Developers can now run llama.cpp models faster, especially when using draft‑based generation that relies on multi‑token prediction.

## Sources & references

1. [ggml-org/llama.cpp b11474](https://github.com/ggml-org/llama.cpp/releases/tag/b11474) – llama.cpp, 2026-10-07

Last updated: 2026-10-07
