AINewsnow

R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]

Here's my *first* implementation of KVA projectors on QFN (just the uncensored model for now) the highlights are basically as follows for using the projectors at each different layer: Starting at layer 12, prompt processing speeds up 1.85x [1700 t/s -> 3150 t/s] at the tradeoff of increasing perple…

Read the full story at r/LocalLLaMA ↗

Timeline · 3 reports

  1. 2026-09-24 18:45 · r/LocalLLM
    Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s
  2. 2026-09-24 15:53 · r/artificial
    I switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro
  3. 2026-09-24 14:02 · r/LocalLLaMA
    R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]

More stories

  1. Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080 — r/LocalLLaMA
  2. Anyone? — r/ChatGPT
  3. My local 27B model made a complete picture book, checked its own image text, fixed a bad page, and exported the PDF — r/LocalLLM
  4. Benchmarking became easy — r/AI_Agents
  5. DeepSeek is CRAZY — Matthew Berman
  6. Retopologizing facial mesh with Codex + Astra + Blender — r/OpenAI
  7. ChatGPT Pro Max 🤖, Muse realtime avatar 🎭, DeepSeek $1B ARR 💰 — TLDR AI
  8. What I've Learned About DeepSeek Harness — KDnuggets

Get the daily brief of stories like this at 6:30 every morning →