Oossa

llama.cpp adds CUDA kernels for block widths over 512

The Oct 8 2026 release adds new fast‑Fourier‑Hadamard kernels that handle widths from 1,024 to 8,192 on NVIDIA GPUs.

NoteBy Published by Oossa: 1 min read

The llama.cpp project released version b11482 on Oct 8 2026. It adds CUDA FWHT kernels that work for block widths above 512, extending support to 1,024‑8,192 for both FP32 and FP16 data. The change reuses the existing F16 infrastructure and passes tests on an A10 GPU. Existing warp‑width kernels (64‑512) remain unchanged.

Why it matters

GPU users can now run larger‑width matrix multiplications faster with llama.cpp.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.