# Memory layer lets AI reuse 50‑million‑token context faster and cheaper

> The galahad‑kv package stores KV state on encrypted NVMe, cutting compute time 2.8‑4.3× and energy 8.8‑12.3× for long texts.

Oossa · 2026-10-09 · https://oossa.com/en/memory-layer-lets-ai-reuse-50-million-token-context-faster-and-cheaper

Researchers released a public memory layer called galahad‑kv that saves a language model’s key‑value state to encrypted local NVMe disk. It lets the model reload a 16,000‑token block without recomputing it. Tests on a single NVIDIA H100 with Gemma 4 12B and 31B models showed loading was 2.8‑4.3 times faster and used 8.8‑12.3 times less GPU energy over a 50 million‑token stream. The 31B model answered factual questions from earlier in the text correctly 98 % of the time.

## The facts

- Memory layer stores KV state for 16,000‑token blocks on encrypted NVMe.
- Loading a block was 2.8×‑4.3× faster and used 8.8×‑12.3× less GPU energy.

## Why it matters

It lets developers run very long‑context applications on a single GPU, reducing cost and power use.

## Sources & references

1. [Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute](https://arxiv.org/abs/2610.10845) – arXiv, 2026-10-09

Last updated: 2026-10-09
