Note · 1 min read
Study finds weak agreement from a production SQL judge
Researchers tested an AI judge used in a text-to-SQL pipeline and found it often disagreed with human reviewers. A self-hosted Qwen model performed much better at a fraction of the per-call cost, the paper reports.
Oossa · About Oossa
A team reporting on arXiv found that its deployed GPT-4o mini judge agreed with human reviewers only weakly when checking whether AI-generated SQL matched a question. On a set chosen to include many disagreements, agreement was near zero; the judge also incorrectly flagged 77.1% of cases humans judged faithful.
The researchers say a self-hosted Qwen3.6-27B model reached agreement similar to Claude Opus 4.7, while costing about one three-hundredth as much per call. They caution that the direct comparison involved only 96 examples.
Why it matters
Teams using AI to grade generated SQL may need to measure judge accuracy against human reviews before trusting its decisions.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline ↗ | arXiv | Sep 28, 2026 | arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human a |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.