The ggml‑org team released llama.cpp version b11324. The update adds a memory‑map mode that avoids making a second full‑size copy of each tensor when loading models. The change is available in new binary packages for macOS (Apple Silicon and Intel), iOS, and several Ubuntu builds. The release also notes contributions from Claude and NVIDIA engineer Pranesh Gonegandla.
Why it matters
Developers can run larger language models on the same hardware, which helps hobbyists and small teams with limited memory.