AINewsnow

67-84 t/s DeepSeek flash v4 off 2x GX10s

This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.

Finally achieved usable results with 2 gx10 at over 65 tokens a second sustained. The 2570 prompt eval is really crucial for me as well. Overall stoked 10/10 edit: I followed this setup with 2 ASUS GX10 DGX computers :) https://github.com/tonyd2wild/DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-08-29 18:59 · r/LocalLLaMA
    67-84 t/s DeepSeek flash v4 off 2x GX10s

More stories

  1. Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal
  2. DeepSeek’s Insane New Architecture — Two Minute Papers
  3. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  4. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  5. Ai used for chatbots — r/artificial
  6. I gave 6 different AIs the same 5 questions — r/AI_Agents
  7. PromptDeck v1.1.0 – open-source desktop app to benchmark local AND cloud LLMs side-by-side (Ollama, LM Studio + OpenRouter, Groq, DeepSeek…) — r/LocalLLaMA
  8. What’s your favorite AI model for coding right now? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →