The ggml‑org team released llama.cpp version b11471. It adds a new set of 30090 plamo fim tokens to the model vocabulary. The release also provides ready‑to‑run binaries for macOS Apple Silicon, macOS Intel, iOS, and several Linux builds (CPU, Vulkan, CUDA 12.8). Users can download the appropriate package from the GitHub release page.
Why it matters
Developers can now run llama.cpp locally on more devices without building from source, saving time and simplifying testing.