AINewsnow

X2 r9700 running 3.8 27B MXFP4 at 710 t/s decode aggregate @10 streams

This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.

As the title says...due to the diligent work by members of the Launch80 AMD discord server, Qwen 3.8 27B now can run INSANELY fast. Image has all the details, but feel free to ask any questions.

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-08-31 23:07 · r/LocalLLM
    X2 r9700 running 3.8 27B MXFP4 at 710 t/s decode aggregate @10 streams

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. Post-training image models for fandom — Character.AI Blog
  3. Testing Qwen 3.8 27B running locally on a single 5090 — r/LocalLLM
  4. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  6. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  7. Qwen Developers on X: "Qwen-Image 2.1 is going open source" — r/StableDiffusion
  8. Optimizing DGX with Qwen 3.8 Flash Next (open to other models!) — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →