AINewsnow

2× Tesla P100 (2016 cards) in 2026: 110 tok/s on a 30B MoE, 16 tok/s at 1M context

I tested two air cooled P100s with llama.cpp. 545 runs across 29 models from 2B to 122B, at context lengths from empty to 1M tokens. Highlights: 110 tok/s: Nemotron-3.5-Lightning-30B-A3B Q4_0 writing code with its MTP draft head. It still does 39 tok/s at 256k context and 16 tok/s at 1M. MoE models…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-23 05:21 · r/LocalLLM
    2× Tesla P100 (2016 cards) in 2026: 110 tok/s on a 30B MoE, 16 tok/s at 1M context

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. Transformers now runs llama.cpp quants — Hugging Face Blog
  3. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  4. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  5. I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
  6. Dual B60 24GB Performance — r/LocalLLM
  7. My contribution to the local AI community: 9 abliterated models, 99 GGUF quantizations in progress — r/huggingface
  8. Performance tune for gemma4-26b-a4b flash attention shape. by frobnitzem · Pull Request #28450 · ggml-org/llama.cpp · GitHub — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →