Qwen 3.8 (27B + Flash Next) on a 128GB Strix Halo laptop as a Claude Opus replacement for agentic coding. AA 40 vs 42, 10-15 tok/s decode, 3 min cold prefill
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: I run Qwen 3.8 (27B and Flash Next) on a 128GB Strix Halo laptop for most of my coding now. It can replace Opus 4.6 to 4.8 for agentic coding if you dont mind a task taking 2 or 3 times longer. Setup: ASUS ROG Flow Z13, Ryzen AI Max+ 395, 128GB unified memory, Arch Linux. llama.cpp as backen…
Read the full story at r/LocalLLM ↗
Timeline · 4 reports
- 2026-09-14 04:30 · r/LocalLLaMA
Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra - 2026-09-12 21:48 · r/huggingface
qwen 3.8 27b ridge m4 24gb. Which models are you using for agentic coding? - 2026-09-12 21:08 · r/LocalLLaMA
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo - 2026-09-11 05:45 · r/LocalLLM
Qwen 3.8 (27B + Flash Next) on a 128GB Strix Halo laptop as a Claude Opus replacement for agentic coding. AA 40 vs 42, 10-15 tok/s decode, 3 min cold prefill