AINewsnow

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the age…

Read the full story at r/MachineLearning ↗

Timeline · 1 report

  1. 2026-10-06 07:03 · r/MachineLearning
    SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

More stories

  1. Mistral releases Mistral Large 4, dubbed "le Chonk", a 1T-parameter open-weight model for general agentic capabilities, trained on 4,000 Grace Blackwell GPUs (Sabrina Ortiz/The Deep View) — Techmeme
  2. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  3. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  4. OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
  5. can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
  6. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  7. Introducing Mistral Large 4 — Mistral AI News
  8. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →