AINewsnow

Qwen 3.8 (27B + Flash Next) on a 128GB Strix Halo laptop as a Claude Opus replacement for agentic coding. AA 40 vs 42, 10-15 tok/s decode, 3 min cold prefill

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

TL;DR: I run Qwen 3.8 (27B and Flash Next) on a 128GB Strix Halo laptop for most of my coding now. It can replace Opus 4.6 to 4.8 for agentic coding if you dont mind a task taking 2 or 3 times longer. Setup: ASUS ROG Flow Z13, Ryzen AI Max+ 395, 128GB unified memory, Arch Linux. llama.cpp as backen…

Read the full story at r/LocalLLM ↗

Timeline · 4 reports

  1. 2026-09-14 04:30 · r/LocalLLaMA
    Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra
  2. 2026-09-12 21:48 · r/huggingface
    qwen 3.8 27b ridge m4 24gb. Which models are you using for agentic coding?
  3. 2026-09-12 21:08 · r/LocalLLaMA
    Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
  4. 2026-09-11 05:45 · r/LocalLLM
    Qwen 3.8 (27B + Flash Next) on a 128GB Strix Halo laptop as a Claude Opus replacement for agentic coding. AA 40 vs 42, 10-15 tok/s decode, 3 min cold prefill

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  3. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  4. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  5. I gave 6 different AIs the same 5 questions — r/AI_Agents
  6. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  7. Is it just me or does Qwen 2.1 look like a heavily distilled GTP image version? — r/StableDiffusion
  8. Using local/non-Anthropic LLMs in Claude Desktop on Windows? — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →