# Soft‑prompt tuning helps LLM benchmarks focus on knowledge, not format

> Aleph Alpha shows that training just ten tiny vectors for a minute lets base models answer in the right format, giving clearer scores.

Oossa · 2026-09-29 · https://oossa.com/en/soft-prompt-tuning-helps-llm-benchmarks-focus-on-knowledge-not-format

Aleph Alpha released a method that tweaks only a handful of learnable vectors—called soft prompts—to teach a frozen language model the exact answer format a benchmark expects. The tuning takes about 80 steps, roughly 1.5 minutes for a 7 billion‑parameter model, and improves format‑following accuracy to near‑100 percent without changing the model’s knowledge.

Because the model’s knowledge score stays flat while formatting improves, researchers can compare base models fairly and predict how they will rank after full post‑training, saving weeks of compute.

## The facts

- 10 soft‑prompt vectors represent 0.0006 % of a 7 B model’s parameters
- Training takes ~80 steps, about 1.5 minutes on a 7 B model

## Why it matters

It lets developers evaluate and rank models quickly without costly full training, focusing on what the model actually knows.

## Sources & references

1. [Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation](https://www.aleph-alpha.com/en/blog/soft-prompt-tuning-for-fair-and-efficient-llm-benchmark-evaluation/) – Aleph Alpha, 2026-09-29

Last updated: 2026-09-29
