# Self‑pruning transformer cuts KV‑cache size up to 25×

> Researchers introduce Universal Attention, a trainable method that adaptively removes low‑impact tokens, achieving 10× compression on typical tasks and 25× on 16k‑token sequences.

Oossa · 2026-10-08 · https://oossa.com/en/self-pruning-transformer-cuts-kv-cache-size-up-to-25

A team led by Ahan Gupta presented a new architecture called Universal Attention. It keeps the standard RoPE positional embeddings and Softmax attention but adds a trainable decay mechanism that decides which tokens to drop during inference. The method prunes the key‑value cache—the memory used by large language models—without hurting accuracy. Tests show a ten‑fold reduction in cache size on normal data and a twenty‑five‑fold reduction when processing 16,000‑token inputs, while still improving downstream performance.

## The facts

- 10× KV‑cache compression on natural‑language tasks
- 25× compression at 16k token length

## Why it matters

Smaller cache needs mean cheaper, faster deployment of large language models on limited‑memory devices.

## Sources & references

1. [A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention](https://arxiv.org/abs/2610.09051) – arXiv cs.LG, 2026-10-08

Last updated: 2026-10-08
