# llama.cpp fixes CPU handling for BF16 matrix inputs

> Release b11292 adds CPU support for a BF16 matrix input used by depthwise convolution and tightens Vulkan’s support checks.

Oossa · 2026-09-30 · https://oossa.com/en/llama-cpp-fixes-cpu-handling-for-bf16-matrix-inputs

Produced and translated with AI assistance. Check the original sources below.

The llama.cpp b11292 release fixes a CPU backend limitation affecting depthwise 1D convolution when its kernel uses BF16, a 16-bit number format. The CPU backend now widens that input to 32-bit float for the calculation, matching the arithmetic used by the Metal matrix-vector kernel.

The release also changes Vulkan’s checks: it accepts BF16 as the second matrix input only when the first input is BF16 too. Other combinations are marked unsupported so the scheduler keeps them on the CPU. The notes describe added tests, but do not say when the changes will reach a packaged app or which users will receive them.

## The facts

- The release is llama.cpp b11292, published September 30, 2026.
- The notes cover a depthwise convolution test across F32, F16 and BF16 kernels, plus three matrix-multiplication cases with BF16 as the second input.

## Why it matters

For developers running BF16 depthwise convolution, the CPU backend can now handle a combination it previously rejected; the notes do not specify when downstream apps will include the fix.

## Sources & references

1. [ggml-org/llama.cpp b11292](https://github.com/ggml-org/llama.cpp/releases/tag/b11292) – llama.cpp, 2026-09-30

Last updated: 2026-09-30
