[Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD)…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-06 13:18 · r/LocalLLaMA
[Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models