OossaAI is evolving fast. We explain it simply.
Newsletter

Note · 1 min read

Study finds weak agreement from a production SQL judge

Researchers tested an AI judge used in a text-to-SQL pipeline and found it often disagreed with human reviewers. A self-hosted Qwen model performed much better at a fraction of the per-call cost, the paper reports.

Oossa · About Oossa

A team reporting on arXiv found that its deployed GPT-4o mini judge agreed with human reviewers only weakly when checking whether AI-generated SQL matched a question. On a set chosen to include many disagreements, agreement was near zero; the judge also incorrectly flagged 77.1% of cases humans judged faithful.

The researchers say a self-hosted Qwen3.6-27B model reached agreement similar to Claude Opus 4.7, while costing about one three-hundredth as much per call. They caution that the direct comparison involved only 96 examples.

Why it matters

Teams using AI to grade generated SQL may need to measure judge accuracy against human reviews before trusting its decisions.

Was this article useful?

Sources & references

#SourceOutletDateKey takeaway
1Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline ↗arXivSep 28, 2026arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human a

1 sources

Last updated:

Oossallms.txt.md

Share

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Study finds weak agreement from a production SQL judge – Oossa