AINewsnow

Mercury released Mercury Decide, I benchmarked it.

Again, I am not English native, expect typo. Let me start with Accuracy, main score of benchmark. Jev 83.3% Solar Decide 57.8% Kev 72.2% Mercury Decide 66.7% But this is no ordinary benchmark, this uses uncontaminated, Korean focused dataset. Dataset & More information It is focused to benchmarking…

Read the full story at r/ArtificialInteligence ↗

Timeline · 1 report

  1. 2026-10-01 05:29 · r/ArtificialInteligence
    Mercury released Mercury Decide, I benchmarked it.

More stories

  1. Ollama now supports Jev-style decision models — Ollama Blog
  2. I rebuilt a Jev-style classifier on Qwen3.5-4B: shared-prefix tree, open weights, fine-tunable, ~140 ms on one H100 — r/learnmachinelearning
  3. Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights) — r/LocalLLaMA
  4. New frontier AI models, TypeSafe’s Jev AI, & NASA’s IBM collab — Mixture of Experts (IBM)
  5. OpenAI's Jev clone could help the frontier lab stop its swarming agents — TechCrunch AI
  6. LLM or JEV? Why not both? - introducing a hybrid Gemma4 approach — r/LocalLLM
  7. NIRNAY: a 450M open model that beats Jev — DEV Community — Machine Learning
  8. Laya/JEV play Magic the Gathering — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →