How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 920k tokens kv cache
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
BetterBench Decode Results On the R9700's I figured out the best path and quality was to get W4A8 running. AMD had also just dropped their AWQ MXFP4 quant of Qwen3.8 27B which is what I am running along with FP8 kv cache. The quality has been great, I ran comparisons across a mini SWE bench and in…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-01 22:35 · r/LocalLLaMA
How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache - 2026-09-01 02:35 · r/LocalLLM
How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 920k tokens kv cache