It doesn’t make sense to me that Gemini 3.8 beats all the models on Terminal-Bench 2.1, yet comes in last on Terminal-Bench 4.0.
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
Coverage of "It doesn’t make sense to me that Gemini 3.8 beats all the models on Terminal-Bench 2.1, yet comes in last on Terminal-Bench 4.0." from 1 source, with a live timeline of who reported what and when.
Read the full story at r/GeminiAI ↗