Is there any way to speed up prefill?
.. I am using M4 pro macbook pro with 48gb of ram I am running qwen3.8 32b with ollama I am using it with VSCode and tried with both Copilot and Continue extensions What I noticed is that the prefill takes very long time (~7-10minutes) After this time the model is definitely usable. After a bit of…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-28 08:17 · r/LocalLLM
Is there any way to speed up prefill?