Oossa

llama.cpp adds BF16/FP16 support in tinyBLAS for x86 CPUs

The open‑source llama.cpp library now vectorizes BF16, FP16 and FP32 tails in its tinyBLAS backend, improving CPU inference speed.

NoteOossaPublished by Oossa: 1 min read

The ggml‑org team released llama.cpp version b11398. It adds support for BF16, FP16 and FP32 tail processing in the tinyBLAS CPU backend on x86 platforms. The changes also vectorize those tails for faster computation. Pre‑built binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu architectures are provided. The update includes test tweaks to ensure accurate comparisons when using reference implementations.

Why it matters

Developers can now run LLM inference faster on standard CPUs without GPU acceleration.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.