Oossa

llama.cpp fixes CPU handling for BF16 matrix inputs

Release b11292 adds CPU support for a BF16 matrix input used by depthwise convolution and tightens Vulkan’s support checks.

NoteOossaPublished by Oossa: 1 min read

The llama.cpp b11292 release fixes a CPU backend limitation affecting depthwise 1D convolution when its kernel uses BF16, a 16-bit number format. The CPU backend now widens that input to 32-bit float for the calculation, matching the arithmetic used by the Metal matrix-vector kernel.

The release also changes Vulkan’s checks: it accepts BF16 as the second matrix input only when the first input is BF16 too. Other combinations are marked unsupported so the scheduler keeps them on the CPU. The notes describe added tests, but do not say when the changes will reach a packaged app or which users will receive them.

Why it matters

For developers running BF16 depthwise convolution, the CPU backend can now handle a combination it previously rejected; the notes do not specify when downstream apps will include the fix.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.