Nemotron 3 Super (120B) at 43 tok/s on one RTX 4090, 2.5× faster than llama.cpp
We got NVIDIA's Nemotron 3 Super (120B, 12B active) running on a single consumer GPU at ~3 bits per weight. GPU tok/s RTX 5090 62 RTX 4090 43 RTX 3090 37 RTX 4080 SUPER 34 32 GB RAM PC 13–18 Same 4090, same model: llama.cpp + Unsloth Q2_K_XL does 17 tok/s. Our file is smaller too (48.7 GB vs 54.7 G…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-11 00:35 · r/LocalLLM
Nemotron 3 Super (120B) at 43 tok/s on one RTX 4090, 2.5× faster than llama.cpp