AINewsnow

Six extraction models against 97 hand-labeled cells: the last misses were ASR mangling model names, not the LLM

I have a small pipeline that takes videos where people pit LLMs against each other on test tasks. It transcribes the audio, then has an LLM pull out structured records of which model passed which test. I wanted to know which model could do the extraction step best. Setup The gold set is 6 videos wi…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-30 04:23 · r/AI_Agents
    Six extraction models against 97 hand-labeled cells: the last misses were ASR mangling model names, not the LLM

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  3. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  4. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  5. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  6. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  7. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  8. OpenAI launches Dots, its Muse competitor — The Verge AI

Get the daily brief of stories like this at 6:30 every morning →