Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K
I ran with sparse attention on for the better part of a month before I sat down and benchmarked it at short context, and it had been costing me time that whole stretch without me noticing. There are two separate things in the engine that both get called sparse attention, and only one of them is the…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-10-02 07:24 · r/machinelearningnews
Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K