AINewsnow

NInfer with improved prefix caching and tool call fixes

I made a custom fork of NInfer for the RTX5090 on Windows which replaces the prefix caching system with one that works really well and also fixes various other issues including tool calling. On an example agentic coding workload my fork reduces TTFT by an average of 80% and increases the cache hit…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-26 07:18 · r/LocalLLM
    NInfer with improved prefix caching and tool call fixes

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  4. GPT‑6 Sol and Luna: Cheaper, but Worse Where It Matters — r/OpenAI
  5. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  6. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  7. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  8. Appeals Court Lets the Pentagon Designate Anthropic a Supply-Chain Risk — Wired AI

Get the daily brief of stories like this at 6:30 every morning →