Researchers released a public memory layer called galahad‑kv that saves a language model’s key‑value state to encrypted local NVMe disk. It lets the model reload a 16,000‑token block without recomputing it. Tests on a single NVIDIA H100 with Gemma 4 12B and 31B models showed loading was 2.8‑4.3 times faster and used 8.8‑12.3 times less GPU energy over a 50 million‑token stream. The 31B model answered factual questions from earlier in the text correctly 98 % of the time.
Why it matters
It lets developers run very long‑context applications on a single GPU, reducing cost and power use.