AINewsnow

Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM

Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15…

Read the full story at r/ArtificialInteligence ↗

Timeline · 6 reports

  1. 2026-10-06 05:39 · r/LocalLLaMA
    unsloth/Qwen3.8-Flash-Next-GGUF is being updated
  2. 2026-10-06 01:47 · r/LocalLLaMA
    Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra
  3. 2026-10-05 23:48 · r/LocalLLM
    RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to
  4. 2026-10-05 14:40 · r/LocalLLM
    Uncensored models: what they are and how they work, with a file-by-file hash check of one uncensored Qwen3.8-Flash-Next upload against Qwen's original
  5. 2026-10-04 20:23 · r/LocalLLM
    Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram
  6. 2026-10-03 08:33 · r/ArtificialInteligence
    Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM

More stories

  1. can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
  2. SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE — r/LocalLLaMA
  3. Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset — r/LocalLLaMA
  4. Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp — r/LocalLLaMA
  5. From one telegram message to a comic book then deploy to web site —— totally local and free — r/LocalLLM
  6. Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen — r/LocalLLaMA
  7. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost
  8. One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui

Get the daily brief of stories like this at 6:30 every morning →