Exa AI announced ATLAS, a benchmark that measures the accuracy and completeness of AI agents that rely on web search. It uses 547 real‑world research tasks drawn from anonymized search demand. The benchmark shows that even the most expensive agents miss about a third of the correct results and that no system costing under $1 per task reaches a row F1 score above 0.5.
Why it matters
Developers can use ATLAS to identify which search backends and agent configurations give the best return for money, guiding more reliable AI‑driven research tools.