# LRCC adds token‑dependent computation to low‑rank compressed LLMs

> The new LRCC method lets Llama and Qwen models choose cheaper or richer pathways per token, boosting accuracy without extra hardware.

Oossa · 2026-10-08 · https://oossa.com/en/lrcc-adds-token-dependent-computation-to-low-rank-compressed-llms

Researchers introduced Low‑Rank Conditional Computation (LRCC), a way to make compressed language models smarter. Each Transformer block gets a tiny router that picks among several low‑rank paths based on the current token. The routers are the only part trained; the low‑rank matrices stay frozen. Tested on Llama‑2‑7B, Llama‑3.2‑1B and Qwen, LRCC gave a 7.6‑point jump in downstream accuracy on Llama‑2‑7B compared with static low‑rank compression, and better perplexity at the same decoding speed.

## The facts

- 7.6 percentage‑point gain in average downstream accuracy on Llama‑2‑7B
- Paper posted on 2026‑10‑08

## Why it matters

Developers can run cheaper, token‑adaptive models that stay accurate, saving compute costs for applications that use Llama or Qwen.

## Sources & references

1. [LRCC: Generalizing Low-Rank Compression with Conditional Computation](https://arxiv.org/abs/2610.08858) – arXiv, 2026-10-08

Last updated: 2026-10-08
