AINewsnow

The computer-use benchmark where Astra passes 2.8 percent of programmatic tests

This story is from 2026-09-21. It is preserved in the archive; the latest stories are on the live feed.

Most computer-use benchmarks test one surface at a time. An agent gets a web browser, a command-line terminal, or an operating system desktop, with a discrete target like filling out a form or editing a config file. Real engineering tasks rarely stay inside one boundary. A developer looks at a runn…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-21 16:16 · DEV Community — AI
    The computer-use benchmark where Astra passes 2.8 percent of programmatic tests

More stories

  1. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  2. Bessent hails US-China AI dialogue ahead of Trump-Xi meeting — Financial Times AI
  3. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  4. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  5. Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters — The Decoder
  6. Amazon blocks Meta’s Muse AI agent — The Verge AI
  7. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  8. AI hallucination of Chinese nuclear components almost led to US military attack — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →