AINewsnow

Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15…

Read the full story at r/LocalLLaMA ↗

Timeline · 5 reports

  1. 2026-10-04 20:23 · r/LocalLLM
    Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram
  2. 2026-10-03 20:15 · r/LocalLLM
    Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
  3. 2026-10-03 20:06 · r/LocalLLaMA
    Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
  4. 2026-10-03 08:33 · r/ArtificialInteligence
    Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM
  5. 2026-10-03 06:49 · r/LocalLLaMA
    Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

More stories

  1. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Claude and Grok built me a local monitoring setup for my AI box: two dashboards, one for the machine, one for model training — r/LocalLLM
  3. Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
  4. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  5. From one telegram message to a comic book then deploy to web site —— totally local and free — r/LocalLLM
  6. Local model for Hermes on a 48GB M5 Pro Mac mini that’s also my daily dev machine? — r/LocalLLM
  7. Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell — PyTorch Blog
  8. One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui

Get the daily brief of stories like this at 6:30 every morning →