LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.03079v1 Announce Type: new Abstract: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv cs.LG
LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference