The llama.cpp b11292 release fixes a CPU backend limitation affecting depthwise 1D convolution when its kernel uses BF16, a 16-bit number format. The CPU backend now widens that input to 32-bit float for the calculation, matching the arithmetic used by the Metal matrix-vector kernel.
The release also changes Vulkan’s checks: it accepts BF16 as the second matrix input only when the first input is BF16 too. Other combinations are marked unsupported so the scheduler keeps them on the CPU. The notes describe added tests, but do not say when the changes will reach a packaged app or which users will receive them.
Why it matters
For developers running BF16 depthwise convolution, the CPU backend can now handle a combination it previously rejected; the notes do not specify when downstream apps will include the fix.