NInfer RTX 3090 Prefill Optimization: 32% Faster at 32K Context (785 to 1,037 tok/s)
I spent a few days repeatedly asking ChatGPT/Codex to find another optimization for NInfer’s RTX 3090 path. Most ideas made things slower and were discarded, but the useful changes added up. With Qwen3.8-27B and the same 32K token input benchmark setup, prefill performance improved from about 785 t…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-24 11:15 · r/LocalLLM
NInfer RTX 3090 Prefill Optimization: 32% Faster at 32K Context (785 to 1,037 tok/s)