Qwen3.8 flash next on 12gb vram + 32 gb ram rest offloaded to ssd
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
According to this video: https://m.youtube.com/watch?v=IH8XmxiwliQ&pp=0gcJCSQMAYcqIYzv&ra=m its possible to get 24 token per second when running qwen3.8 flash next with lot of it offloaded to the ssd. He also said to set threat count to 6 instead of 12 (in his case) which boosted performance from 1…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-17 22:10 · r/LocalLLM
Qwen3.8 flash next on 12gb vram + 32 gb ram rest offloaded to ssd