ETA on Qwen4-35b using GPU+RAM+NVMe?
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Qwen-Flash-Next uses incredible qwen4 architecture and according to YT videos runs 20tps+ on 16gb GPU due to offloading on RAM+NVMe. How long till we get an absolute MONSTER 35b Qwen4 model that smokes 3.8-27b and runs on 16gb VRAM at 40 tps and 128k context? Anyone hearing anything or seen leaks?…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-14 19:10 · r/LocalLLM
ETA on Qwen4-35b using GPU+RAM+NVMe?