AINewsnow

Why real-time ASR is so expensive to serve — and how to fix it

Hey everyone, we built .wave, an inference engine for models that run continuously. We’re now using it to serve NVIDIA Nemotron 3.5 ASR Streaming 0.6B at $0.00045/minute . I wanted to share how the engine works, why it matters for streaming workloads, and what we measured. When you’re building a vo…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-10-03 16:19 · r/AI_Agents
    Why real-time ASR is so expensive to serve — and how to fix it

More stories

  1. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  2. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  3. NVIDIA DGX Spark 64GB. The Reason behind current Spark shortages. — r/LocalLLM
  4. The Nvidia Shield TV Is 7 Years Old. It Just Got a $100 Price Hike — Wired AI
  5. Nvidia Partner Hon Hai Beats Sales Estimates Due to AI Frenzy — Bloomberg AI
  6. Nvidia’s Valuations Show AI Rally Isn’t a Bubble, DBS Says — Bloomberg AI
  7. Scoop: A powerful new model from startup Reflection is set to shake up the AI race — Axios AI+
  8. Why tech companies are racing to put AI data centers in space — Fast Company AI

Get the daily brief of stories like this at 6:30 every morning →