# llama.cpp updates add GLM‑5.3‑Flash support and performance tweaks

> The open‑source LLM runtime llama.cpp v0.1.0 adds GLM‑5.3‑Flash model support and several speed and memory fixes.

Oossa · 2026-09-30 · https://oossa.com/en/llama-cpp-updates-add-glm-5-3-flash-support-and-performance-tweaks

Produced and translated with AI assistance. Check the original sources below.

The llama.cpp project released a new version on September 30, 2026. It adds support for GLM‑5.3‑Flash (also called GLM5‑Next), a newer language model architecture. The update also reduces compute buffer size, speeds up long‑context decoding, and fixes several bugs around token handling and GPU memory. The changes are aimed at developers who run LLMs locally on CPUs or GPUs.

## The facts

- Release date: 2026‑09‑30 ("Wed Sep 30 2026")
- Adds GLM‑5.3‑Flash (GLM5‑Next) model support

## Why it matters

The new features let users run the latest GLM models faster and with less memory on their own hardware.

## Sources & references

1. [ggml-org/llama.cpp b11279](https://github.com/ggml-org/llama.cpp/releases/tag/b11279) – llama.cpp, 2026-09-30
2. [add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp](https://www.reddit.com/r/LocalLLaMA/comments/1wu0bdf/add_glm53flash_glm5next_support_by_timkhronos/) – Reddit r/LocalLLaMA, 2026-09-30
3. [add GLM-5.3-Flash (GLM5-Next) support (#27773) · ggml-org/llama.cpp@649dcb1](https://www.reddit.com/r/LocalLLaMA/comments/1wu3wxu/add_glm53flash_glm5next_support_27773/) – Reddit r/LocalLLaMA, 2026-09-30

Last updated: 2026-09-30
