Same model, same laptop, one app is way slower. What do you check first?
Say the same weights answer in seconds in one app and take ages in another. Even "hi" is slow. Would you start with the prompt the app actually sends, the chat template, or tool calls? I dont want to swap models before finding out what extra stuff the app is doing. Anyone chased this down? submitte…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-29 19:14 · r/LocalLLM
Same model, same laptop, one app is way slower. What do you check first?