For dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost
For the last few months, I ran DeepSeek v4.0 flash (NVFP4). First 0731 , then visionexp because it was a free improvement. I got around 65 tps decode and almost 2k prefill, and ran 4-5 agents in parallel, totalling around 200 tps cumulative decode. Because of this, I did not feel like switching to…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-04 21:53 · r/LocalLLaMA
For dual DGX spark users; GLM 5.3 flash got a 50%+ performance boost