Oossa

llama.cpp update stops overwriting busy slots

The Oct 10, 2026 release fixes a bug where a new request could corrupt an ongoing generation by rewriting its prompt cache.

NoteBy Published by Oossa: 1 min read

The llama.cpp project released version b11552 on 2026‑10‑10. The change stops a request from updating the prompt cache of a slot that is already busy. Previously, the busy slot would be overwritten, causing the generation to continue with the wrong context. Now the busy slot is returned unchanged and the new request waits until the slot is free.

Why it matters

Developers using llama.cpp will see fewer generation errors when multiple requests share the same GPU memory.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.