# llama.cpp adds BF16/FP16 support in tinyBLAS for x86 CPUs

> The open‑source llama.cpp library now vectorizes BF16, FP16 and FP32 tails in its tinyBLAS backend, improving CPU inference speed.

Oossa · 2026-10-04 · https://oossa.com/en/llama-cpp-adds-bf16-fp16-support-in-tinyblas-for-x86-cpus

The ggml‑org team released llama.cpp version b11398. It adds support for BF16, FP16 and FP32 tail processing in the tinyBLAS CPU backend on x86 platforms. The changes also vectorize those tails for faster computation. Pre‑built binaries for macOS (Apple Silicon and Intel), iOS, and several Ubuntu architectures are provided. The update includes test tweaks to ensure accurate comparisons when using reference implementations.

## The facts

- Release version b11398 published on Oct 4, 2026
- Adds BF16/FP16/FP32 tail support in tinyBLAS for x86

## Why it matters

Developers can now run LLM inference faster on standard CPUs without GPU acceleration.

## Sources & references

1. [ggml-org/llama.cpp b11398](https://github.com/ggml-org/llama.cpp/releases/tag/b11398) – llama.cpp, 2026-10-04

Last updated: 2026-10-04
