Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
arXiv:2609.30290v1 Announce Type: new Abstract: Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv cs.CL
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline