AINewsnow

i made a lot of unofficial tests for different 3 and 4 bit quants of qwen 3.8-27b on my local work on rtx 3090 ti with 96gb ram, and ThinkingCap-Qwen3.6-27B is way better and faster than qwen 3.8-27b, and glm 5.3 and muse spark 1.2, so for me ai benchmarks are useless

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

RTX 3090 Ti 24 GB · 96gb ram - Windows · llama.cpp- DeepSeek Harness ngl 99 -c %CTX% -fa on -np 1 -ctk q8_0 -ctv q8_0 -temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 / --min-p 0.05 (to test if it will fix it) --presence-penalty 0.0 --repeat-penalty 1.0 (+MTP) - ---------------------------------------…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-05 09:51 · r/LocalLLM
    I ran Qwen 3.8 27B on my new MacBook Pro M5 Max and on my RTX 5090 workstation. The Mac held up way better than I expected.
  2. 2026-09-03 23:08 · r/LocalLLM
    i made a lot of unofficial tests for different 3 and 4 bit quants of qwen 3.8-27b on my local work on rtx 3090 ti with 96gb ram, and ThinkingCap-Qwen3.6-27B is way better and faster than qwen 3.8-27b, and glm 5.3 and muse spark 1.2, so for me ai benchmarks are useless

More stories

  1. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  4. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  5. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  6. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/ArtificialInteligence
  7. OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news — ThursdAI
  8. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →