AINewsnow

Nemotron 3 Super (120B) at 43 tok/s on one RTX 4090, 2.5× faster than llama.cpp

We got NVIDIA's Nemotron 3 Super (120B, 12B active) running on a single consumer GPU at ~3 bits per weight. GPU tok/s RTX 5090 62 RTX 4090 43 RTX 3090 37 RTX 4080 SUPER 34 32 GB RAM PC 13–18 Same 4090, same model: llama.cpp + Unsloth Q2_K_XL does 17 tok/s. Our file is smaller too (48.7 GB vs 54.7 G…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-11 00:35 · r/LocalLLM
    Nemotron 3 Super (120B) at 43 tok/s on one RTX 4090, 2.5× faster than llama.cpp

More stories

  1. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  2. I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp — r/LocalLLaMA
  3. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  4. Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects — r/LocalLLaMA
  5. Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s — r/LocalLLaMA
  6. Qwen3.8-Flash-Next (125B) at ~100 tok/s on an M5 Ultra Mac Studio with llama.cpp — r/LocalLLM
  7. Success with Qwen3.8 27B GSQ-RCO-IQ3_S on 16GB VRAM — r/LocalLLM
  8. Created laya : Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →