Six extraction models against 97 hand-labeled cells: the last misses were ASR mangling model names, not the LLM
I have a small pipeline that takes videos where people pit LLMs against each other on test tasks. It transcribes the audio, then has an LLM pull out structured records of which model passed which test. I wanted to know which model could do the extraction step best. Setup The gold set is 6 videos wi…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-30 04:23 · r/AI_Agents
Six extraction models against 97 hand-labeled cells: the last misses were ASR mangling model names, not the LLM