28 TPS on Qwen2.5-7B across two separate cloud regions over public WAN using speculative decoding + CUDA Graphs [P]
This story is from 2026-08-23. It is preserved in the archive; the latest stories are on the live feed.
been building ShardFlow for the past few months, a distributed LLM inference framework that splits any HuggingFace transformer across N GPU machines and uses neural speculative decoding to deal with WAN latency. the setup for the benchmark: two T4 nodes in separate GCP regions (Iowa + Oregon) talki…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-08-23 12:30 · r/MachineLearning
28 TPS on Qwen2.5-7B across two separate cloud regions over public WAN using speculative decoding + CUDA Graphs [P]