AINewsnow

Same model, same laptop, one app is way slower. What do you check first?

Say the same weights answer in seconds in one app and take ages in another. Even "hi" is slow. Would you start with the prompt the app actually sends, the chat template, or tool calls? I dont want to swap models before finding out what extra stuff the app is doing. Anyone chased this down? submitte…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 19:14 · r/LocalLLM
    Same model, same laptop, one app is way slower. What do you check first?

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  3. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  4. OpenAI launches Dots, its Muse competitor — The Verge AI
  5. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  6. Anthropic warns of ‘existential risks to humanity’ in IPO prospectus — Financial Times AI
  7. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  8. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →