AINewsnow

Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

TL;DR: Halogen v0.14.0 with its native .hgn weights is the fastest, followed by gufo and CIRU. Halogen is closed source and runs in Docker. gufo is open source and loads 4x faster from cold. gufo is also fastest to first token on follow-ups (1.6-1.9 s against 2.6-3.3 s). I benchmarked different eng…

Read the full story at r/LocalLLM ↗

Timeline · 4 reports

  1. 2026-10-01 05:27 · r/LocalLLaMA
    Qwen Flash Next MTP work restarted
  2. 2026-09-30 17:15 · r/LocalLLaMA
    How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?
  3. 2026-09-30 14:29 · r/LocalLLM
    Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
  4. 2026-09-30 09:41 · r/LocalLLM
    Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

More stories

  1. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
  3. Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
  4. What local AI model is good for game decomps/recomps? — r/LocalLLaMA
  5. Which LLM is best for coding/agents if I have dual R9700 GPUs? — r/LocalLLM
  6. Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s — r/LocalLLM
  7. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
  8. What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →