The llama.cpp project released a new version on September 30, 2026. It adds support for GLM‑5.3‑Flash (also called GLM5‑Next), a newer language model architecture. The update also reduces compute buffer size, speeds up long‑context decoding, and fixes several bugs around token handling and GPU memory. The changes are aimed at developers who run LLMs locally on CPUs or GPUs.
Why it matters
The new features let users run the latest GLM models faster and with less memory on their own hardware.