# llama.cpp update cuts indexer memory use in half

> The b11372 release changes the Qwen4exp indexer to use far less RAM, with no speed loss, for long‑context runs.

Oossa · 2026-10-03 · https://oossa.com/en/llama-cpp-update-cuts-indexer-memory-use-in-half

The ggml‑org/llama.cpp project released version b11372 on Oct 3 2026. The update rewrites the Qwen4exp indexer so each head computes its score separately and writes results in place. This halves the memory needed for the indexer score buffers, which were the biggest memory users in long‑context graphs. The change does not affect compute speed and also adds support for the new lightning indexer on CUDA, Metal and Vulkan back‑ends.

## The facts

- Release tag b11372 published on Sat Oct 03 2026
- Indexer score memory reduced by about 50 %

## Why it matters

Developers running large language models with long context will be able to fit bigger prompts on the same hardware.

## Sources & references

1. [ggml-org/llama.cpp b11372](https://github.com/ggml-org/llama.cpp/releases/tag/b11372) – llama.cpp, 2026-10-03

Last updated: 2026-10-03
