Oossa

llama.cpp release adds memory‑map option to cut tensor copies

The open‑source LLM runtime now lets users map model data directly, saving RAM on macOS, iOS and Linux builds.

NoteOossaPublished by Oossa: 1 min read

The ggml‑org team released llama.cpp version b11324. The update adds a memory‑map mode that avoids making a second full‑size copy of each tensor when loading models. The change is available in new binary packages for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds. The release also notes contributions from Claude and NVIDIA engineer Pranesh Gonegandla.

Why it matters

Developers can run larger language models on the same hardware, which helps hobbyists and small teams with limited memory.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.