Oossa

AI agent teams cost more but rarely improve results

Vals AI’s study shows multi‑agent setups can spend up to five times the tokens of a single model for only tiny quality gains.

By Published by Oossa: 1 min read

Event date:

Marvin Meyer · Unsplash

A new evaluation by Vals AI compared single AI agents to teams of agents on its Vibe Code Bench. The test used two leading models – GPT‑6 Sol and Claude Opus 5.5 – in solo mode and in groups of up to 100 agents. Teams consumed between 1.8 × and 5.1 × more tokens than a lone model, but delivered almost no better scores.

What the numbers show

Out of four head‑to‑head comparisons, only one was statistically significant: GPT‑6 Sol at a medium reasoning level scored 7.3 points higher when run as a team. At the maximum reasoning setting, neither model gained a measurable edge. The data suggests that the extra compute cost of agent swarms rarely translates into higher quality, especially when the models are already running at full capacity.

Why speed‑up doesn’t equal better output

OpenAI researchers echo the finding. In a podcast, Noam Brown noted that adding agents mainly speeds up tasks that parallelize well, such as web research or math, but does not improve the answer itself. Four agents solved a task twice as fast but also cost twice as much. Scaling to dozens of agents showed diminishing returns, and for creative work like novel writing, large swarms add no value.

Why it matters

For developers budgeting cloud compute, the study warns that adding many AI agents will likely raise expenses without improving the product’s quality. Companies planning to speed up routine tasks should weigh the token cost against the modest speed gains, and may be better off optimizing single‑agent prompts.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.