Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work
While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to…
Read the full story at r/singularity ↗
Timeline · 1 report
- 2026-09-30 20:44 · r/singularity
Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work