RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to
I got two big MoE models running locally on a budget AMD setup and wanted to share the numbers and steps for other RX 7600 owners. Everything is measured on my own machine. Rig: RX 7600 8 GB (gfx1102), Ryzen 5 5600 (AVX2, no AVX-512), 64 GB DDR4, B450 board (PCIe 3.0 x8), Ubuntu 26.04, system ROCm…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-10-06 05:39 · r/LocalLLaMA
unsloth/Qwen3.8-Flash-Next-GGUF is being updated - 2026-10-05 23:48 · r/LocalLLM
RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to