Oossa

Exa launches ATLAS benchmark for agentic web search

The new ATLAS benchmark grades how well AI agents retrieve complete, accurate results from real web searches, highlighting current cost and performance gaps.

NoteBy Published by Oossa: 1 min read

Exa AI announced ATLAS, a benchmark that measures the accuracy and completeness of AI agents that rely on web search. It uses 547 real‑world research tasks drawn from anonymized search demand. The benchmark shows that even the most expensive agents miss about a third of the correct results and that no system costing under $1 per task reaches a row F1 score above 0.5.

Why it matters

Developers can use ATLAS to identify which search backends and agent configurations give the best return for money, guiding more reliable AI‑driven research tools.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.