The ggml‑org team posted llama.cpp version b11270 on Sep 30, 2026. It tweaks a GPU operation used by AMD’s HIP platform, replacing a slower byte‑subtraction instruction with a faster one. The release also bundles fresh binary packages for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds, including Vulkan‑enabled GPU versions.
Why it matters
The GPU tweak can speed up inference on AMD hardware, and the ready‑to‑run binaries make it easier for developers to try llama.cpp on more platforms.