Jev: one judgment call, or twelve dimension scores? I measured both on three classification tasks
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
The obvious way to use Jev (TypeSafe AI's System One judgment API) is one question per row, then threshold the score. I wanted to know whether the other shape earns its extra tokens: 12 to 14 narrow questions per row, cached scores, weights fitted locally on my own labels. Three tasks, both shapes,…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-17 16:13 · DEV Community — Machine Learning
Jev: one judgment call, or twelve dimension scores? I measured both on three classification tasks