The ggml‑org team released llama.cpp version b11262 on Sep 29, 2026. The update adds AVX512‑FP16 support, which does half‑precision (f16) dot products but accumulates them in single‑precision (f32) for better accuracy. It is a low‑level change that speeds up inference on CPUs that have the AVX512‑FP16 instruction set. The release also bundles new binaries for macOS, iOS, and several Linux builds.
Why it matters
It lets developers run LLMs faster on modern CPUs without losing numerical quality.