AINewsnow

NInfer RTX 3090 Prefill Optimization: 32% Faster at 32K Context (785 to 1,037 tok/s)

I spent a few days repeatedly asking ChatGPT/Codex to find another optimization for NInfer’s RTX 3090 path. Most ideas made things slower and were discarded, but the useful changes added up. With Qwen3.8-27B and the same 32K token input benchmark setup, prefill performance improved from about 785 t…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-24 11:15 · r/LocalLLM
    NInfer RTX 3090 Prefill Optimization: 32% Faster at 32K Context (785 to 1,037 tok/s)

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  3. OpenAI Agent Hacked Australian Government Website — Wall Street Journal Technology
  4. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  5. Anthropic and OpenAI roll out cheaper models in first release since call for slowdown — CNBC Technology
  6. Anthropic launches Claude Opus 5.5, promising Fable-level performance at a lower price — Mashable AI
  7. Help? — r/GeminiAI
  8. Ringg’s AI agents resolve up to 65% of customer calls with OpenAI — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →