Distilling Sequential Computation in Transformer Language Models
arXiv:2609.27233v1 Announce Type: new Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their r…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-24 04:00 · arXiv cs.CL
Distilling Sequential Computation in Transformer Language Models