What's 2 dgx spark real decode with context filled near full and concurrency
TL;DR: Considering two DGX Sparks for 5+ simultaneous coding/security research agents, running something like DeepSeek V4 Flash at a good-quality quant. What decode tokens/sec per agent can I realistically expect at 100–140K occupied context each? Like shared 1M but maybe 256k for each Most benchma…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-06 11:41 · r/LocalLLM
What's 2 dgx spark real decode with context filled near full and concurrency