AINewsnow

It doesn’t make sense to me that Gemini 3.8 beats all the models on Terminal-Bench 2.1, yet comes in last on Terminal-Bench 4.0.

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

Coverage of "It doesn’t make sense to me that Gemini 3.8 beats all the models on Terminal-Bench 2.1, yet comes in last on Terminal-Bench 4.0." from 1 source, with a live timeline of who reported what and when.

Read the full story at r/GeminiAI ↗

Timeline · 1 report

  1. 2026-09-02 17:11 · r/GeminiAI
    It doesn’t make sense to me that Gemini 3.8 beats all the models on Terminal-Bench 2.1, yet comes in last on Terminal-Bench 4.0.

More stories

  1. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  2. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  3. AI skills — r/AI_Agents
  4. Gemini 4 Pro vs Fable 5 vs GPT6 Astra — r/GeminiAI
  5. AI models are not hacking “autonomously” — r/artificial
  6. Plugin4Shell and NIST IR 8587, days apart: what actually authorizes an AI agent’s action? — r/AI_Agents
  7. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  8. Gemini 2.5 pro model disappeared in AI Studio — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →