Troubleshooting errors when running models on a phone
So, here’s a breakdown of the issues: There was a race condition between the text-to-speech (TTS) processing and the model's token generation; they were loading simultaneously, causing the generation speed to drop from 15 tokens/sec to 5 tokens/sec. I won't be fixing this right now-lots of cool voi…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-04 05:32 · r/LocalLLM
Troubleshooting errors when running models on a phone