Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions
This story is from 2026-09-29. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.31675v1 Announce Type: new Abstract: We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world gen…
Read the full story at arXiv cs.LG ↗
Timeline · 2 reports
- 2026-09-29 04:00 · arXiv cs.AI
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks - 2026-09-29 04:00 · arXiv cs.LG
Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions