The ggml‑org team released llama.cpp version b11293. It adds BF16 (a half‑precision number format) unary, GLU, binary and scale operations for both CPU and CUDA. The change also updates CUDA kernels to fix a BF16 build issue. The update is available for many platforms, from macOS and Windows to Linux and Android.
Why it matters
Developers can now run llama.cpp models faster on hardware that supports BF16, such as newer GPUs.