AINewsnow

DeepSeek V4 Flash and Qwen 3.8 Flash: Single RTX Pro 6000

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

qwen 3.8 flash next(nvfp4 + offloading): 4 concurrent sessions at 256k: 45 t/ts each, ttft: 6s 8 concurrent sessions at 256k: 35 t/s each, ttft: 6.4s 16 concurrent sessions at 256k: 30 t/s each, ttft: 12s Deepseek v4 flash(native with offloading): 8 concurrent sessions at 256k: 18 t/s each, ttft: 9…

Read the full story at r/LocalLLM ↗

Timeline · 3 reports

  1. 2026-09-07 02:31 · r/LocalLLM
    Unsloth template for DeepSeek-V4-Flash-Vision-Exp-GGUF
  2. 2026-09-06 20:15 · r/LocalLLaMA
    DeepSeek-V4-Flash-Vision Q8 vs Qwen3.8-Flash-Next Q8
  3. 2026-09-04 22:34 · r/LocalLLM
    DeepSeek V4 Flash and Qwen 3.8 Flash: Single RTX Pro 6000

More stories

  1. I gave 6 different AIs the same 5 questions — r/AI_Agents
  2. Alibaba's Qwen3.8-Omni-Flash Cuts Video AI Costs by 89% With Agent Tool Use — AlphaSignal
  3. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  4. Alibaba releases Qwen-Image-2.1, a 7B open-weight model it says outperforms most closed-source models, with native transparency and up to ten reference images (Qwen) — Techmeme
  5. Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal
  6. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  7. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  8. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →