Oossa

Memory layer lets AI reuse 50‑million‑token context faster and cheaper

The galahad‑kv package stores KV state on encrypted NVMe, cutting compute time 2.8‑4.3× and energy 8.8‑12.3× for long texts.

NoteBy Published by Oossa: 1 min read

Researchers released a public memory layer called galahad‑kv that saves a language model’s key‑value state to encrypted local NVMe disk. It lets the model reload a 16,000‑token block without recomputing it. Tests on a single NVIDIA H100 with Gemma 4 12B and 31B models showed loading was 2.8‑4.3 times faster and used 8.8‑12.3 times less GPU energy over a 50 million‑token stream. The 31B model answered factual questions from earlier in the text correctly 98 % of the time.

Why it matters

It lets developers run very long‑context applications on a single GPU, reducing cost and power use.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.