AINewsnow

DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)

I've been running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700s (32 GB each, 192 GB system RAM) using affinity( https://codeberg.org/StillDeadcode/affinity ), an inference engine written by **Yoshi Exeler (StillDeadcode)** specifically for DeepSeek-V4-Flash on one or two RDNA4 cards. All the c…

Read the full story at r/LocalLLaMA ↗

Timeline · 4 reports

  1. 2026-09-24 18:45 · r/LocalLLM
    Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s
  2. 2026-09-24 15:53 · r/artificial
    I switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro
  3. 2026-09-24 14:02 · r/LocalLLaMA
    R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]
  4. 2026-09-23 17:49 · r/LocalLLaMA
    DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)

More stories

  1. DeepSeek details DSec sandbox infrastructure for agent training — TechNode
  2. [AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M — Latent Space
  3. Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact — r/LocalLLaMA
  4. Jev's calibration was measured. The LLMs won [D] — r/MachineLearning
  5. My local 27B model made a complete picture book, checked its own image text, fixed a bad page, and exported the PDF — r/LocalLLM
  6. Perplexity Brings Portable Computer to AMD Ryzen AI Max PCs — AlphaSignal
  7. What I've Learned About DeepSeek Harness — KDnuggets
  8. Google invented the Transformer, so why does using Gemini still feel like a chore? — r/ArtificialInteligence

Get the daily brief of stories like this at 6:30 every morning →