The ggml‑org team released llama.cpp version b11489. The update speeds up Q6_K weight dequantisation – a step that converts compressed model weights back to usable numbers – and doubles the unrolling factor, according to the release notes. Pre‑built binaries are now available for Apple Silicon, Intel macOS, iOS, and several Ubuntu configurations including Vulkan and CUDA builds.
Why it matters
The faster dequantisation means developers can run LLMs on the same hardware with lower latency.