AINewsnow

RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to

I got two big MoE models running locally on a budget AMD setup and wanted to share the numbers and steps for other RX 7600 owners. Everything is measured on my own machine. Rig: RX 7600 8 GB (gfx1102), Ryzen 5 5600 (AVX2, no AVX-512), 64 GB DDR4, B450 board (PCIe 3.0 x8), Ubuntu 26.04, system ROCm…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-10-06 05:39 · r/LocalLLaMA
    unsloth/Qwen3.8-Flash-Next-GGUF is being updated
  2. 2026-10-05 23:48 · r/LocalLLM
    RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to

More stories

  1. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Looking for developer-friendly inference providers who give you enough API credits to experiment [D] — r/MachineLearning
  3. Qwen-Image-2.1-Multiple-Angles-LoRA — r/StableDiffusion
  4. Local AI ecosystem overview — r/LocalLLaMA
  5. Strata: Qwen 3.8 Flash next in loop — r/LocalLLM
  6. Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature — r/LocalLLaMA
  7. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  8. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →