What benchmark do you wish someone would build? 👀
Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today. Have you ever wanted to e…
Read the full story at r/AI_Agents ↗
Timeline · 2 reports
- 2026-10-02 08:43 · r/learnmachinelearning
What benchmark do you wish someone would build? - 2026-10-02 08:36 · r/AI_Agents
What benchmark do you wish someone would build? 👀