AINewsnow

49 tok/s 4060ti

Getting 40-50 toks/s qwen3.8 unsloth iq4 xs with ollama off of a 4060ti 16gb. Context is 32k but it doesn’t seem an issue with autocompact in vscode. Using claude to manage and direct qwen. Code reviews are clean and it’s fast enough for my workflow. Qwen is saving me money as I have gone from $300…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-25 01:43 · r/LocalLLM
    49 tok/s 4060ti

More stories

  1. Can we take a moment to appreciate that with 950 Claude agents running for only 21 hours searching genomic data, Anthropic may have found a new CRISPR-like gene-editing mechanism — r/singularity
  2. AI 週報 — 2026-09-18 to 2026-09-25 | 模型晶片雙線交火,前沿落地訊號開始分流 — DEV Community — Machine Learning
  3. We audited 30 coding agent test edits and found an 86.7% false alarm rate in naive test tampering detection. Here is what we learned/ — r/PromptEngineering
  4. NInfer Qwen 3.8-27B uncensored on RTX 5090 175 tok/s changed my life — r/LocalLLM
  5. I got Mimo 2.6-Flash-RL running at 30-45 tok/s on the Strix Halo 128GB 2TB — r/LocalLLM
  6. 5070 ti + 64 ram Qwen 3.8 27b — r/LocalLLM
  7. I built a system where Claude Code delegates heavy coding work to a local Qwen model — built a full roguelite overnight without burning through my Pro quota — r/LocalLLM
  8. Qwen 3.8 27B on one 5090: 22 GB VRAM, 175k context — a coding driver, not a Claude replacement — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →