Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s
I’ve been building a local AI server around four unlocked CMP 170HX 64GB cards and finally have DeepSeek V4 Flash Vision stable across all four GPUs. Hardware 4× NVIDIA CMP 170HX 64GB 256GB aggregate HBM2e Runtime Ubuntu 24.04 vLLM Pipeline parallel = 4 max_model_len = 262144 FP8 KV cache DSpark sp…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-26 17:41 · r/LocalLLaMA
85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri - 2026-09-24 18:45 · r/LocalLLM
Benchmarking DeepSeek V4 Flash on 4× CMP 170HX 64GB: 256GB HBM, PP4, 262K context, up to ~95 tok/s