# Study finds weak agreement from a production SQL judge

> Researchers tested an AI judge used in a text-to-SQL pipeline and found it often disagreed with human reviewers. A self-hosted Qwen model performed much better at a fraction of the per-call cost, the paper reports.

Oossa · 2026-09-29 · https://oossa.com/en/study-finds-weak-agreement-from-a-production-sql-judge

A team reporting on arXiv found that its deployed GPT-4o mini judge agreed with human reviewers only weakly when checking whether AI-generated SQL matched a question. On a set chosen to include many disagreements, agreement was near zero; the judge also incorrectly flagged 77.1% of cases humans judged faithful.

The researchers say a self-hosted Qwen3.6-27B model reached agreement similar to Claude Opus 4.7, while costing about one three-hundredth as much per call. They caution that the direct comparison involved only 96 examples.

## The facts

- GPT-4o mini’s Cohen’s kappa was 0.04 on a disagreement-enriched sample and 0.42 on a random spot-check.
- Qwen3.6-27B reached kappa 0.72; Claude Opus 4.7 reached 0.71.

## Why it matters

Teams using AI to grade generated SQL may need to measure judge accuracy against human reviews before trusting its decisions.

## Sources & references

1. [Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline](https://arxiv.org/abs/2609.30290) – arXiv, 2026-09-28

Last updated: 2026-09-29
