A Terminal-Bench task frontier agents could not pass
This story is from 2026-10-07. It is preserved in the archive; the latest stories are on the live feed.
I wrote a task for Terminal-Bench, a public benchmark of hard command-line tasks for AI agents. The agent has to build a redactor for court-filing PDFs: black out every visible name and ID on a list, leave nothing recoverable in the file, and change nothing else on the page. I ran it three times ea…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-07 17:46 · DEV Community — AI
A Terminal-Bench task frontier agents could not pass
More stories
- Introducing Mistral Large 4 — Mistral AI News
- Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
- GPT-6 and Intelligent UI for everyone — OpenAI News
- Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
- Sharing AI progress in mathematics — OpenAI News
- NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
- Surface RTX Spark Dev Box is available for preorder for $5,999 — The Verge AI
- Introducing Playground: Create and play custom games — Google AI Blog
Get the daily brief of stories like this at 6:30 every morning →