NInfer with improved prefix caching and tool call fixes
I made a custom fork of NInfer for the RTX5090 on Windows which replaces the prefix caching system with one that works really well and also fixes various other issues including tool calling. On an example agentic coding workload my fork reduces TTFT by an average of 80% and increases the cache hit…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-26 07:18 · r/LocalLLM
NInfer with improved prefix caching and tool call fixes