Oossa

llama.cpp adds WebGPU support for duplicate operation

A fix in llama.cpp version b11523 lets WebGPU handle the DUP op, preventing crashes when running mixed‑batch models.

NoteBy Published by Oossa: 1 min read

The open‑source llama.cpp library released version b11523 on Oct 9, 2026. The update fixes a crash that occurred when the library tried to use the DUP operation on WebGPU. Previously, DUP ran on the CPU while other work ran on WebGPU, leading to a buffer mismatch and a crash. DUP now follows the same code path as the CPY (copy) and CONT ops, so it stays on WebGPU.

Why it matters

WebGPU users can run llama.cpp models with mixed batches without the program crashing.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.