AINewsnow

Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2). One-command install, everything pinned.

Another Qwen 3.8 on a R9700 post, but I thought I'd share my repo for anyone who might benefit. I might be wrong, but I don't think I have seen anyone with it this fast on windows. I've spent the last few weeks tuning Qwen3.8-27B on one AMD Radeon AI PRO R9700 (32 GB, RDNA4), and I've packaged the…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-08 10:18 · r/LocalLLM
    Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2). One-command install, everything pinned.

More stories

  1. Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance — r/LocalLLM
  2. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA
  3. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Uncensored models: what they are and how they work, with a file-by-file hash check of one uncensored Qwen3.8-Flash-Next upload against Qwen's original — r/LocalLLM
  5. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM
  6. Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
  7. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  8. Qwen 3.8 27B with Asymmetric Dual GPUs (RTX 3080 20GB Mod + RTX 3070): Pipeline vs Tensor Parallelism, Oculink — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →