2× Tesla P100 (2016 cards) in 2026: 110 tok/s on a 30B MoE, 16 tok/s at 1M context
I tested two air cooled P100s with llama.cpp. 545 runs across 29 models from 2B to 122B, at context lengths from empty to 1M tokens. Highlights: 110 tok/s: Nemotron-3.5-Lightning-30B-A3B Q4_0 writing code with its MTP draft head. It still does 39 tok/s at 256k context and 16 tok/s at 1M. MoE models…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-23 05:21 · r/LocalLLM
2× Tesla P100 (2016 cards) in 2026: 110 tok/s on a 30B MoE, 16 tok/s at 1M context