i stopped caring which model scores higher and started counting how often each one lies to me
benchmarks tell me nothing about my actual day. so for two weeks i kept a tally every time gemini and one other model said something wrong in an area i know cold. the scores nobody publishes are the ones that decide which one i open. gemini was better on factual recall for my niche, worse on admitt…
Read the full story at r/GeminiAI ↗
Timeline · 1 report
- 2026-10-01 13:04 · r/GeminiAI
i stopped caring which model scores higher and started counting how often each one lies to me
More stories
- OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
- OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns — New York Times Technology
- Google's first Gemini 4 model is 'Argon' — Engadget
- Gemini 4 Argon: our next era of frontier intelligence — Google DeepMind Blog
- Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times Technology
- Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
- Google unveils Gemini 4 Argon with SOTA score on DeepSWE — TestingCatalog AI News
- Gemini 4 Pro 🔪 — r/GeminiAI
Get the daily brief of stories like this at 6:30 every morning →