The llama.cpp project released version b11482 on Oct 8 2026. It adds CUDA FWHT kernels that work for block widths above 512, extending support to 1,024‑8,192 for both FP32 and FP16 data. The change reuses the existing F16 infrastructure and passes tests on an A10 GPU. Existing warp‑width kernels (64‑512) remain unchanged.
Why it matters
GPU users can now run larger‑width matrix multiplications faster with llama.cpp.