# llama.cpp release adds OpenVINO 2026.4.1 support and MoE fixes

> The Oct 3 2026 update upgrades the OpenVINO backend, improves performance and corrects MoE model handling for llama.cpp users.

Oossa · 2026-10-03 · https://oossa.com/en/llama-cpp-release-adds-openvino-2026-4-1-support-and-moe-fixes

The llama.cpp library released a new version (b11374) on Oct 3 2026. It updates the OpenVINO backend to version 2026.4.1, adds new conv‑fusion, fixes GPU buffer errors and corrects axis handling for mixture‑of‑experts (MoE) models. The changes boost token‑generation speed for fused MoE models and prevent crashes on some Intel GPUs.

## The facts

- Release date: Oct 3, 2026 ("Sat Oct 03 2026 13:51:21 GMT+0200")
- OpenVINO upgraded to 2026.4.1; token‑generation speed for gemma‑4‑26B improved from 66 t/s to 1 609 t/s when fused

## Why it matters

Developers running llama.cpp on Intel GPUs can now use OpenVINO with higher speed and fewer crashes.

## Sources & references

1. [ggml-org/llama.cpp b11374](https://github.com/ggml-org/llama.cpp/releases/tag/b11374) – llama.cpp, 2026-10-03
2. [ggml-org/llama.cpp b11377](https://github.com/ggml-org/llama.cpp/releases/tag/b11377) – llama.cpp, 2026-10-03

Last updated: 2026-10-03
