A team led by Ahan Gupta presented a new architecture called Universal Attention. It keeps the standard RoPE positional embeddings and Softmax attention but adds a trainable decay mechanism that decides which tokens to drop during inference. The method prunes the key‑value cache—the memory used by large language models—without hurting accuracy. Tests show a ten‑fold reduction in cache size on normal data and a twenty‑five‑fold reduction when processing 16,000‑token inputs, while still improving downstream performance.
Why it matters
Smaller cache needs mean cheaper, faster deployment of large language models on limited‑memory devices.