Reduced my Jev judge’s calibration error [D]
I benchmarked Jev on TRIVIA+ dataset using an untouched 645-example test set. Before calibration: ECE: 0.0982 After learning from human-labelled examples: ECE: 0.0313 That’s a 68.1% reduction in calibration error . But hallucination-detection F1 only moved: 0.5833 → 0.5877 So what improved? Not the…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-28 02:26 · r/MachineLearning
Reduced my Jev judge’s calibration error [D]