MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
This story is from 2026-09-22. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing syst…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.AI
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators