The llama.cpp library released version b11559 on Oct 11, 2026. It adds OpenCL improvements for faster factor‑allocation and supports Gemma‑4 decoding on Qualcomm E4B and E2B GPUs. The update also optimizes DK64 and DK128 decode paths for various quantisation settings. The changes were co‑authored by Hongqiang Wang of Qualcomm.
Why it matters
Developers can now run Gemma‑4 models faster on compatible Qualcomm GPUs, expanding edge‑AI capabilities.