Why real-time ASR is so expensive to serve — and how to fix it
Hey everyone, we built .wave, an inference engine for models that run continuously. We’re now using it to serve NVIDIA Nemotron 3.5 ASR Streaming 0.6B at $0.00045/minute . I wanted to share how the engine works, why it matters for streaming workloads, and what we measured. When you’re building a vo…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-10-03 16:19 · r/AI_Agents
Why real-time ASR is so expensive to serve — and how to fix it