# llama.cpp update stops overwriting busy slots

> The Oct 10, 2026 release fixes a bug where a new request could corrupt an ongoing generation by rewriting its prompt cache.

Oossa · 2026-10-10 · https://oossa.com/en/llama-cpp-update-stops-overwriting-busy-slots

The llama.cpp project released version b11552 on 2026‑10‑10. The change stops a request from updating the prompt cache of a slot that is already busy. Previously, the busy slot would be overwritten, causing the generation to continue with the wrong context. Now the busy slot is returned unchanged and the new request waits until the slot is free.

## The facts

- Release version b11552 published on Sat Oct 10 2026
- Fix prevents prompt cache updates on busy id_slot

## Why it matters

Developers using llama.cpp will see fewer generation errors when multiple requests share the same GPU memory.

## Sources & references

1. [ggml-org/llama.cpp b11552](https://github.com/ggml-org/llama.cpp/releases/tag/b11552) – llama.cpp, 2026-10-10

Last updated: 2026-10-10
