Benchmarked 14 open models, mostly decision models, on five decisions from our own products; a naive Bayes baseline kept up on the hardest
TL;DR: We asked 14 open models, zero-shot, five choose-an-option questions from our own products. On the hardest (which of six characters said a line), a naive Bayes trained on 227 labelled lines kept up with the best small decision models (68 of 108, against 66 to 72; no paired test told them apar…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-10-01 04:17 · r/machinelearningnews
Benchmarked 14 open models, mostly decision models, on five decisions from our own products; a naive Bayes baseline kept up on the hardest
More stories
- Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
- Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
- OpenAI pauses AI model training after another agent bypasses network restrictions — InfoWorld AI
- Introducing dots — OpenAI News
- Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
- FTC launches broad investigation into Anthropic, OpenAI — Washington Post AI
- Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
- Gemini 4 Argon has a 1M-token output limit, up from 64K for prior models, and initially costs $2/1M input and $10/1M output tokens, rising to $4 and $20 later (Matthias Bastian/The Decoder) — Techmeme
Get the daily brief of stories like this at 6:30 every morning →