AINewsnow

From benchmarking to fine-tuning: what I learned about small decision models

This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.

I tested two small decision models. First I tested TypeSafe AI's hosted model, Jev. Then I tested Laya, an independent open-source alternative. I started out benchmarking and ended up fine-tuning. Here is what I found, in the order I found it. Part 1: Benchmarking Jev TypeSafe AI launched Jev on Se…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-30 05:47 · DEV Community — AI
    From benchmarking to fine-tuning: what I learned about small decision models

More stories

  1. Ollama now supports Jev-style decision models — Ollama Blog
  2. Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights) — r/LocalLLaMA
  3. OpenDecider: distilling calibrated "System One" decision models from open teachers, evaluated head-to-head against Laya and TypeSafe's Jev — r/machinelearningnews
  4. New frontier AI models, TypeSafe’s Jev AI, & NASA’s IBM collab — Mixture of Experts (IBM)
  5. LLM or JEV? Why not both? - introducing a hybrid Gemma4 approach — r/LocalLLM
  6. How long until OpenAI release their version of Jev? — r/AI_Agents
  7. autotrust/JEV-27B-VL: first decision model that learned to see without a single image of training — r/LocalLLM
  8. More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with Jev — arXiv cs.AI

Get the daily brief of stories like this at 6:30 every morning →