Oossa

llama.cpp update cuts indexer memory use in half

The b11372 release changes the Qwen4exp indexer to use far less RAM, with no speed loss, for long‑context runs.

NoteOossaPublished by Oossa: 1 min read

The ggml‑org/llama.cpp project released version b11372 on Oct 3 2026. The update rewrites the Qwen4exp indexer so each head computes its score separately and writes results in place. This halves the memory needed for the indexer score buffers, which were the biggest memory users in long‑context graphs. The change does not affect compute speed and also adds support for the new lightning indexer on CUDA, Metal and Vulkan back‑ends.

Why it matters

Developers running large language models with long context will be able to fit bigger prompts on the same hardware.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.