# llama.cpp adds Vulkan duplicate‑row optimizations

> The open‑source LLaMA inference library now skips redundant work when GPU rows repeat, improving speed on Vulkan devices.

Oossa · 2026-10-11 · https://oossa.com/en/llama-cpp-adds-vulkan-duplicate-row-optimizations

The llama.cpp project released version b11554 on 2026‑10‑11. The update adds Vulkan‑specific fixes that detect duplicate row IDs in matrix multiplications and handle them in a pre‑pass instead of looping. It also improves the “hoist row ids” optimization so it always runs and can emit multiple compact tiles. The changes let the GPU launch fewer work‑groups and exit early when possible.

## The facts

- Release version b11554 published on Sun Oct 11 2026
- Vulkan duplicate‑row handling added to llama.cpp

## Why it matters

GPU users of llama.cpp can see faster inference because the library now avoids repeating the same calculations.

## Sources & references

1. [ggml-org/llama.cpp b11554](https://github.com/ggml-org/llama.cpp/releases/tag/b11554) – llama.cpp, 2026-10-11

Last updated: 2026-10-11
