AINewsnow

Ito: streaming speech synthesis in 4.89 MB, verified in ESP32-S3 emulation

This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.

We've released Ito's inference engine, ESP32-S3 firmware and two English voice models. The latest demo lists 4.05M parameters, about 4M. The chip weights occupy 4.89 MB per voice. The engine produces 24 kHz speech using integer arithmetic. What runs where Text becomes phonemes on the host using esp…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-05 09:07 · DEV Community — AI
    Ito: streaming speech synthesis in 4.89 MB, verified in ESP32-S3 emulation

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  3. A model guide for the GPT-6 family — OpenAI News
  4. An OpenAI safety employee has quit and is sounding the alarm — The Verge AI
  5. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  6. Introducing Oscilloscope Diffusion — r/comfyui
  7. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  8. Trump expected to tap DNI Jay Clayton as new AI czar — Axios AI+

Get the daily brief of stories like this at 6:30 every morning →