Shipped a voice agent on the Realtime API. Went through production call logs and found 7 behavioral bugs that no amount of scripted testing would have caught
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Been running a voice agent on the Realtime API in production for a few months, business use case, not consumer-facing. This week I sat down and read through a batch of live call transcripts hunting for weird behavior instead of relying on eval scores. Found a cluster of bugs worth sharing. None of…
Read the full story at r/OpenAI ↗