DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)
I've been running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700s (32 GB each, 192 GB system RAM) using affinity( https://codeberg.org/StillDeadcode/affinity ), an inference engine written by **Yoshi Exeler (StillDeadcode)** specifically for DeepSeek-V4-Flash on one or two RDNA4 cards. All the c…
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-24 18:45 · r/LocalLLM
Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s - 2026-09-24 15:53 · r/artificial
I switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro - 2026-09-24 14:02 · r/LocalLLaMA
R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5] - 2026-09-23 17:49 · r/LocalLLaMA
DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)