Who's getting above 50 tok/s on AMD 9070, R9700 GPUs?
I have this setup: Windows 11 64GB DDR5 RAM @ 6000Mhz 9070XT 16GB 9700 AI Pro 32GB Running Qwen 3.8 Swift Q6 quant with 131K context, cache type K of Q8, V of Q8, with the Dflash2 drafter via llama.cpp Vulkan, my tok/s ceiling appears to be 50 tok/s. **Anyone getting more than this with AMD kit lik…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-21 15:20 · r/LocalLLM
Who's getting above 50 tok/s on AMD 9070, R9700 GPUs?