Grok 4.6 Scores 26% and 88% on the Same Benchmark Line
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
If you have seen a number for Grok 4.6 in the last month, it was probably one of two: about 88%, or about 26%. Both circulate as Terminal-Bench results. Both are real. They are not the same benchmark, and the difference is a version number that most write-ups drop. The primary source is xAI's Grok…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-07 08:01 · DEV Community — AI
Grok 4.6 Scores 26% and 88% on the Same Benchmark Line
More stories
- xAI Ships Grok Voice Transcribe 2.0 With Half the Errors at Same Price — AlphaSignal
- ZCode was allegedly caught uploading workspace/.git records to the cloud. — r/LocalLLaMA
- SpaceXAI's Grok Bot is in early beta - Here's how to try it — Engadget
- The cloud outage that should terrify the CIO — InfoWorld AI
- Ai used for chatbots — r/artificial
- Whats the best and most ethical use case of AI images? — r/ArtificialInteligence
- Funniest AI block yet (Grok) — r/AI_Agents
- I was researching my long-form RP to death — separate story from continuity searches — r/ChatGPTPro
Get the daily brief of stories like this at 6:30 every morning →