AINewsnow

Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions

This story is from 2026-09-29. It is preserved in the archive; the latest stories are on the live feed.

arXiv:2609.31675v1 Announce Type: new Abstract: We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world gen…

Read the full story at arXiv cs.LG ↗

Timeline · 2 reports

  1. 2026-09-29 04:00 · arXiv cs.AI
    EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
  2. 2026-09-29 04:00 · arXiv cs.LG
    Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  4. Anthropic warns of ‘existential risks to humanity’ in IPO prospectus — Financial Times AI
  5. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  6. OpenAI launches Dots, always-on agents powered by GPT-6 Astra with their own cloud computer, in ChatGPT for Pro, Business Premium, and Enterprise users (Rachel Metz/Bloomberg) — Techmeme
  7. OpenAI Scraps Release of New AI Model Over Safety Concerns — Wall Street Journal Technology
  8. OpenAI DevDay 2026: The biggest news and announcements — The Verge AI

Get the daily brief of stories like this at 6:30 every morning →