The llama.cpp project released version b11552 on 2026‑10‑10. The change stops a request from updating the prompt cache of a slot that is already busy. Previously, the busy slot would be overwritten, causing the generation to continue with the wrong context. Now the busy slot is returned unchanged and the new request waits until the slot is free.
Why it matters
Developers using llama.cpp will see fewer generation errors when multiple requests share the same GPU memory.