AINewsnow

Is there any way to speed up prefill?

.. I am using M4 pro macbook pro with 48gb of ram I am running qwen3.8 32b with ollama I am using it with VSCode and tried with both Copilot and Continue extensions What I noticed is that the prefill takes very long time (~7-10minutes) After this time the model is definitely usable. After a bit of…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-28 08:17 · r/LocalLLM
    Is there any way to speed up prefill?

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. Bill Gates warns AI is powerful enough to cause "a billion deaths" — Axios AI+
  3. Scoop: Top AI companies probing tens of thousands of security incidents — Axios AI+
  4. GitHub Copilot app for Beginners: How to build custom workflows with canvases — GitHub Blog
  5. Microsoft releases .NET SDK for AG-UI agent-user interaction protocol — InfoWorld AI
  6. AI agents are the ultimate Aggregators; they reveal apps as a means, not an end, and offering them is tech's biggest prize, with Meta and Microsoft well-poised (Ben Thompson/Stratechery) — Techmeme
  7. Microsoft goes quiet after church groups ask for 1% of data center costs — Ars Technica AI
  8. Ready or not, here come the AI gadgets, from "charms" to smart speakers — CBS News Technology

Get the daily brief of stories like this at 6:30 every morning →