# llama.cpp release adds Q6_K weight dequant speedup

> Version b11489 speeds up Q6_K weight dequantisation and adds extra unrolling, with new binaries for macOS, iOS and Linux.

Oossa · 2026-10-08 · https://oossa.com/en/llama-cpp-release-adds-q6-k-weight-dequant-speedup

The ggml‑org team released llama.cpp version b11489. The update speeds up Q6_K weight dequantisation – a step that converts compressed model weights back to usable numbers – and doubles the unrolling factor, according to the release notes. Pre‑built binaries are now available for Apple Silicon, Intel macOS, iOS, and several Ubuntu configurations including Vulkan and CUDA builds.

## The facts

- Release tag: b11489
- Published: Thu Oct 08 2026

## Why it matters

The faster dequantisation means developers can run LLMs on the same hardware with lower latency.

## Sources & references

1. [ggml-org/llama.cpp b11489](https://github.com/ggml-org/llama.cpp/releases/tag/b11489) – llama.cpp, 2026-10-08

Last updated: 2026-10-08
