The open‑source llama.cpp project released version b11558 on Oct 11, 2026. The update fixes a bug that caused OpenCL kernels to exceed device image limits when using Q4_K quantized weights. It also adds a fallback path for devices that hit those limits, and corrects broadcasting and RMS‑norm calculations for Q4_0 kernels. The changes were co‑authored by Hongqiang Wang of Qualcomm.
Why it matters
Developers using llama.cpp on GPUs can now run Q4_K quantized models on more hardware without crashes.