Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction…
Read the full story at Apple Machine Learning Research ↗
Timeline · 1 report
- 2026-09-24 00:00 · Apple Machine Learning Research
Compressing Streaming Neural Audio Encoders via Latent-Space Distillation