Ito: streaming speech synthesis in 4.89 MB, verified in ESP32-S3 emulation
This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.
We've released Ito's inference engine, ESP32-S3 firmware and two English voice models. The latest demo lists 4.05M parameters, about 4M. The chip weights occupy 4.89 MB per voice. The engine produces 24 kHz speech using integer arithmetic. What runs where Text becomes phonemes on the host using esp…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-05 09:07 · DEV Community — AI
Ito: streaming speech synthesis in 4.89 MB, verified in ESP32-S3 emulation