I tried to bait LLM judges into picking wrong answers. The weak ones rewrote the math to agree.
What I Benchmarked I've spent the last couple of months on a fine-tuning leaderboard where every submission gets scored by an LLM judge. The same model, trained on the same data, could land a 46% win rate on one judged metric and 85% on another. After a while I stopped asking "is my model good?" an…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-11 14:02 · DEV Community — Machine Learning
I tried to bait LLM judges into picking wrong answers. The weak ones rewrote the math to agree.