DeepSeek V4 Flash and Qwen 3.8 Flash: Single RTX Pro 6000
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
qwen 3.8 flash next(nvfp4 + offloading): 4 concurrent sessions at 256k: 45 t/ts each, ttft: 6s 8 concurrent sessions at 256k: 35 t/s each, ttft: 6.4s 16 concurrent sessions at 256k: 30 t/s each, ttft: 12s Deepseek v4 flash(native with offloading): 8 concurrent sessions at 256k: 18 t/s each, ttft: 9…
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-09-07 02:31 · r/LocalLLM
Unsloth template for DeepSeek-V4-Flash-Vision-Exp-GGUF - 2026-09-06 20:15 · r/LocalLLaMA
DeepSeek-V4-Flash-Vision Q8 vs Qwen3.8-Flash-Next Q8 - 2026-09-04 22:34 · r/LocalLLM
DeepSeek V4 Flash and Qwen 3.8 Flash: Single RTX Pro 6000