AINewsnow

Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s

I’ve been building a local AI server around four unlocked CMP 170HX 64GB cards and finally have DeepSeek V4 Flash Vision stable across all four GPUs. Hardware 4× NVIDIA CMP 170HX 64GB 256GB aggregate HBM2e Runtime Ubuntu 24.04 vLLM Pipeline parallel = 4 max_model_len = 262144 FP8 KV cache DSpark sp…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-26 17:41 · r/LocalLLaMA
    85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri
  2. 2026-09-24 18:45 · r/LocalLLM
    Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s

More stories

  1. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  2. Another "Harness matters" post (codex cli > pi and opencode) — r/LocalLLaMA
  3. Worth going from Qwen3.8 27B to flash next or maaybe deepseek v4 flash? — r/LocalLLM
  4. Gemini 3.8 flash VS DeepSeek V4.1 — r/GeminiAI
  5. Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode — r/LocalLLaMA
  6. Which provider actually wins on pure affordability right now for gemma qwen gpt oss and deepseek under one roof — r/AI_Agents
  7. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  8. When is the next generation of "B tier" models releasing? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →