The open‑source llama.cpp library released version b11523 on Oct 9, 2026. The update fixes a crash that occurred when the library tried to use the DUP operation on WebGPU. Previously, DUP ran on the CPU while other work ran on WebGPU, leading to a buffer mismatch and a crash. DUP now follows the same code path as the CPY (copy) and CONT ops, so it stays on WebGPU.
Why it matters
WebGPU users can run llama.cpp models with mixed batches without the program crashing.