AINewsnow

Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to…

Read the full story at r/singularity ↗

Timeline · 1 report

  1. 2026-09-30 20:44 · r/singularity
    Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work

More stories

  1. Gemini 4 Argon: our next era of frontier intelligence — Google DeepMind Blog
  2. Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times AI
  3. Gemini 4 Argon has a 1M-token output limit, up from 64K for prior models; it initially costs $2/1M input and $10/1M output tokens, rising to $4 and $20 later (Matthias Bastian/The Decoder) — Techmeme
  4. Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
  5. Gemini 4 Argon — Hacker News Front Page
  6. Gemini Pro 4 (leak) — r/singularity
  7. Let skills in Gemini tackle your most repetitive tasks — Google Gemini Blog
  8. See what 4 builders are making with Gemini 3.8 Flash — Google Gemini Blog

Get the daily brief of stories like this at 6:30 every morning →