AINewsnow

ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]

Figure 1b from our paper Disclosure: I'm one of the authors (Microsoft). The paper, code, dataset are public and ThinkingBox is on Hugging Face OpenEnv as well. Raw evaluation trajectories are not released Links at the bottom. We wanted to know how much of a single agent success rate survives repet…

Read the full story at r/MachineLearning ↗

Timeline · 1 report

  1. 2026-10-09 00:50 · r/MachineLearning
    ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]

More stories

  1. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  2. Claude Pro vs ChatGPT Plus vs Copilot Premium: which one would you choose for this use case? — r/ChatGPTPro
  3. Microsoft’s new Surface Laptop Ultra finally has a starting price (you should sit down) — ZDNET AI
  4. Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google — r/LocalLLaMA
  5. Moonworks Lunara: Modeling Artistic Intelligence [R] — r/MachineLearning
  6. Rogue AI or human error? The real story behind the OpenAI-Hugging Face incident — Scientific American
  7. OpenAI will watermark ChatGPT outputs by default—but only in the EU — Ars Technica AI
  8. Microsoft releases new Nvidia-chip AI PCs with revamped Windows 11 — TechCrunch AI

Get the daily brief of stories like this at 6:30 every morning →