A new evaluation by Vals AI compared single AI agents to teams of agents on its Vibe Code Bench. The test used two leading models – GPT‑6 Sol and Claude Opus 5.5 – in solo mode and in groups of up to 100 agents. Teams consumed between 1.8 × and 5.1 × more tokens than a lone model, but delivered almost no better scores.
What the numbers show
Out of four head‑to‑head comparisons, only one was statistically significant: GPT‑6 Sol at a medium reasoning level scored 7.3 points higher when run as a team. At the maximum reasoning setting, neither model gained a measurable edge. The data suggests that the extra compute cost of agent swarms rarely translates into higher quality, especially when the models are already running at full capacity.
Why speed‑up doesn’t equal better output
OpenAI researchers echo the finding. In a podcast, Noam Brown noted that adding agents mainly speeds up tasks that parallelize well, such as web research or math, but does not improve the answer itself. Four agents solved a task twice as fast but also cost twice as much. Scaling to dozens of agents showed diminishing returns, and for creative work like novel writing, large swarms add no value.
Why it matters
For developers budgeting cloud compute, the study warns that adding many AI agents will likely raise expenses without improving the product’s quality. Companies planning to speed up routine tasks should weigh the token cost against the modest speed gains, and may be better off optimizing single‑agent prompts.