Researchers introduced Low‑Rank Conditional Computation (LRCC), a way to make compressed language models smarter. Each Transformer block gets a tiny router that picks among several low‑rank paths based on the current token. The routers are the only part trained; the low‑rank matrices stay frozen. Tested on Llama‑2‑7B, Llama‑3.2‑1B and Qwen, LRCC gave a 7.6‑point jump in downstream accuracy on Llama‑2‑7B compared with static low‑rank compression, and better perplexity at the same decoding speed.
Why it matters
Developers can run cheaper, token‑adaptive models that stay accurate, saving compute costs for applications that use Llama or Qwen.