The ggml‑org team released llama.cpp version b11398. It adds support for BF16, FP16 and FP32 tail processing in the tinyBLAS CPU backend on x86 platforms. The changes also vectorize those tails for faster computation. Pre‑built binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu architectures are provided. The update includes test tweaks to ensure accurate comparisons when using reference implementations.
Why it matters
Developers can now run LLM inference faster on standard CPUs without GPU acceleration.