The llama.cpp library released a new version (b11374) on Oct 3 2026. It updates the OpenVINO backend to version 2026.4.1, adds new conv‑fusion, fixes GPU buffer errors and corrects axis handling for mixture‑of‑experts (MoE) models. The changes boost token‑generation speed for fused MoE models and prevent crashes on some Intel GPUs.
Why it matters
Developers running llama.cpp on Intel GPUs can now use OpenVINO with higher speed and fewer crashes.