AINewsnow

anyone experienced this before?

Before doing anything, my models were running at 20tps. After a few hours has gone by + running and ejecting the model, somehow it magically goes down to 4-5tps. I don't have any program installed, the task manager is fine (no heavy programs running). I was only playing around with deepseek harness…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-25 16:28 · r/LocalLLM
    anyone experienced this before?

More stories

  1. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  2. R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5] — r/LocalLLaMA
  3. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  4. When is the next generation of "B tier" models releasing? — r/LocalLLaMA
  5. DeepSeek Elastic Compute (DSec) — Hacker News Front Page
  6. Benchmarking became easy — r/AI_Agents
  7. DeepSeek is CRAZY — Matthew Berman
  8. Retopologizing facial mesh with Codex + Astra + Blender — r/OpenAI

Get the daily brief of stories like this at 6:30 every morning →