Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090
tl;dr: I created a fully-local open-source full-duplex voice agent that rivals GPT-Live on some benchmarks. It uses Voxtral Realtime with a turn-taking head, a microturn-finetuned Gemma 4 12B and Breeze TTS 2 under the hood. Go try it out: https://github.com/speakrail/speakrail Interjections work!…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-05 15:50 · r/LocalLLaMA
Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090