# llama.cpp release adds memory‑map option to cut tensor copies

> The open‑source LLM runtime now lets users map model data directly, saving RAM on macOS, iOS and Linux builds.

Oossa · 2026-10-01 · https://oossa.com/en/llama-cpp-release-adds-memory-map-option-to-cut-tensor-copies

The ggml‑org team released llama.cpp version b11324. The update adds a memory‑map mode that avoids making a second full‑size copy of each tensor when loading models. The change is available in new binary packages for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds. The release also notes contributions from Claude and NVIDIA engineer Pranesh Gonegandla.

## The facts

- Release version b11324 published on 2026-10-01
- Adds llama‑mmap feature to reduce RAM usage

## Why it matters

Developers can run larger language models on the same hardware, which helps hobbyists and small teams with limited memory.

## Sources & references

1. [ggml-org/llama.cpp b11324](https://github.com/ggml-org/llama.cpp/releases/tag/b11324) – llama.cpp, 2026-10-01

Last updated: 2026-10-01
