# llama.cpp adds WebGPU support for duplicate operation

> A fix in llama.cpp version b11523 lets WebGPU handle the DUP op, preventing crashes when running mixed‑batch models.

Oossa · 2026-10-09 · https://oossa.com/en/llama-cpp-adds-webgpu-support-for-duplicate-operation

The open‑source llama.cpp library released version b11523 on Oct 9, 2026. The update fixes a crash that occurred when the library tried to use the DUP operation on WebGPU. Previously, DUP ran on the CPU while other work ran on WebGPU, leading to a buffer mismatch and a crash. DUP now follows the same code path as the CPY (copy) and CONT ops, so it stays on WebGPU.

## The facts

- Version b11523 released on 2026-10-09
- Fix adds WebGPU support for the GGML_OP_DUP operation

## Why it matters

WebGPU users can run llama.cpp models with mixed batches without the program crashing.

## Sources & references

1. [ggml-org/llama.cpp b11523](https://github.com/ggml-org/llama.cpp/releases/tag/b11523) – llama.cpp, 2026-10-09

Last updated: 2026-10-09
