EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
arXiv:2609.31906v1 Announce Type: new Abstract: Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained envi…
Read the full story at arXiv cs.AI ↗
Timeline · 2 reports
- 2026-09-29 04:00 · arXiv cs.LG
Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions - 2026-09-29 04:00 · arXiv cs.AI
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks