# llama.cpp adds CUDA kernels for block widths over 512

> The Oct 8 2026 release adds new fast‑Fourier‑Hadamard kernels that handle widths from 1,024 to 8,192 on NVIDIA GPUs.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-adds-cuda-kernels-for-block-widths-over-512

The llama.cpp project released version b11482 on Oct 8 2026. It adds CUDA FWHT kernels that work for block widths above 512, extending support to 1,024‑8,192 for both FP32 and FP16 data. The change reuses the existing F16 infrastructure and passes tests on an A10 GPU. Existing warp‑width kernels (64‑512) remain unchanged.

## The facts

- Release tag b11482, published Oct 8 2026
- New kernels support widths 1,024‑8,192 for FP32 and FP16

## Why it matters

GPU users can now run larger‑width matrix multiplications faster with llama.cpp.

## Sources & references

1. [ggml-org/llama.cpp b11482](https://github.com/ggml-org/llama.cpp/releases/tag/b11482) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
