The ggml‑org team released llama.cpp version b11487. It adds tiled Q4_K and Q6_K GET_ROWS support for Hexagon processors and improves handling of Q4_K views. Pre‑built binaries are provided for Apple Silicon, Intel macOS, iOS, and several Ubuntu builds (CPU and Vulkan). The update is available for download from the GitHub release page.
Why it matters
Developers can now run more efficient Llama models on Hexagon‑based devices such as Qualcomm chips, improving performance on supported mobile and embedded hardware.