AINewsnow

Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)

My journey: 6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters. 21 tps - that's better, let's see IQ3_XXS 27 tps - nice, but still not my tempo. 51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more.…

Read the full story at r/LocalLLM ↗

Timeline · 4 reports

  1. 2026-10-03 01:47 · r/LocalLLM
    Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context.
  2. 2026-10-01 05:27 · r/LocalLLaMA
    Qwen Flash Next MTP work restarted
  3. 2026-09-30 17:15 · r/LocalLLaMA
    How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?
  4. 2026-09-30 14:29 · r/LocalLLM
    Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)

More stories

  1. Claude and Grok built me a local monitoring setup for my AI box: two dashboards, one for the machine, one for model training — r/LocalLLM
  2. add GLM-5.3-Flash (GLM5-Next) support (#27773) · ggml-org/llama.cpp@649dcb1 — r/LocalLLaMA
  3. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  4. I put Gemma 4 26B and Qwen 3.8 27B (2x3090) in charge of a club in Championship Manager 97/98, against Claude, Grok and DeepSeek. It's running now. — r/LocalLLM
  5. I made a simple gui for claude code, usable with local llms — r/LocalLLM
  6. Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
  7. What local AI model is good for game decomps/recomps? — r/LocalLLaMA
  8. Which LLM is best for coding/agents if I have dual R9700 GPUs? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →