Benchmarking small confidence scoring decision (Jev, Laya) models [P]
I ran my own evals of two small classification models that score confidence over a list of candidate labels instead of generating text: TypeSafe AI's hosted Jev, and Laya, an independent open-source alternative. Some findings that I think apply beyond these two models: 1. Describe labels by what's…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-30 12:41 · r/MachineLearning
Benchmarking small confidence scoring decision (Jev, Laya) models [P]
More stories
- Ollama now supports Jev-style decision models — Ollama Blog
- Amazon releases its own Jev clone as decision models flood the web — TechCrunch AI
- OpenAI answers TypeSafe's Jev with a Decision API built on Luna — The New Stack AI
- I rebuilt a Jev-style classifier on Qwen3.5-4B: shared-prefix tree, open weights, fine-tunable, ~140 ms on one H100 — r/learnmachinelearning
- Vev: Jev-like decision models with vision — 4B/9B, local inference, open weights — r/LocalLLM
- Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights) — r/LocalLLaMA
- New frontier AI models, TypeSafe’s Jev AI, & NASA’s IBM collab — Mixture of Experts (IBM)
- OpenAI's Jev clone could help the frontier lab stop its swarming agents — TechCrunch AI
Get the daily brief of stories like this at 6:30 every morning →