Upgrade existing PC to run Qwen 3.8 Flash Next via Strata or swap to a Strix Halo?
This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.
My daily locally LLM is a Q3 quant of Qwen 3.8:27b, which runs with 140k CTX at circa 40 TPS gen and 500 PP on 24gb across dual GPU (5060ti 16gb + 3060Ti 8gb). I am using Opencode to do single-user agentic coding. It's OK, but even with thinking set to medium, the useful work between initial CTX lo…
Read the full story at r/LocalLLM ↗
Timeline · 9 reports
- 2026-10-07 22:26 · r/LocalLLM
Is anyone running Qwen Flash Next on Strata with q6 or higher? - 2026-10-07 22:02 · r/LocalLLM
Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next - 2026-10-07 20:43 · r/LocalLLaMA
Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B? - 2026-10-06 14:58 · r/LocalLLaMA
Qwen3.8-Flash-Next on Strata - 2026-10-06 10:37 · r/LocalLLaMA
NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s - 2026-10-05 17:38 · r/LocalLLM
Strata: Qwen 3.8 Flash next in loop - 2026-10-05 12:28 · r/LocalLLM
Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance - 2026-10-05 07:27 · r/LocalLLM
4x 3080 20GB (modded, alibaba) + Strata (Qwen Flash Next 125B IQ3_XXS) = 105 t/s generation 5000 prompt processing on 150k context - 2026-10-05 07:21 · r/LocalLLM
Upgrade existing PC to run Qwen 3.8 Flash Next via Strata or swap to a Strix Halo?