ggml‑org released llama.cpp version b11555. The update corrects a bug where AVX2 CPUs did not use the F16C instruction set even when it was available. The release also ships new binary packages for macOS Apple Silicon, macOS Intel, iOS, and several Ubuntu configurations including Vulkan and CUDA 12.8 support.
Why it matters
Developers running Llama models on compatible CPUs will see faster inference without manual patches.