The ggml‑org team released llama.cpp version b11268. The update fixes a bug in the OpenCL driver that handled tensors for the q5_K Adreno GPU kernel. New pre‑built binaries are provided for macOS Apple Silicon, Intel Macs, iOS, and several Linux configurations including CPU, Vulkan and CUDA 12.8.
Why it matters
The fix improves performance and stability for developers running LLMs on Android devices with Adreno GPUs.