i made a lot of unofficial tests for different 3 and 4 bit quants of qwen 3.8-27b on my local work on rtx 3090 ti with 96gb ram, and ThinkingCap-Qwen3.6-27B is way better and faster than qwen 3.8-27b, and glm 5.3 and muse spark 1.2, so for me ai benchmarks are useless
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
RTX 3090 Ti 24 GB · 96gb ram - Windows · llama.cpp- DeepSeek Harness ngl 99 -c %CTX% -fa on -np 1 -ctk q8_0 -ctv q8_0 -temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 / --min-p 0.05 (to test if it will fix it) --presence-penalty 0.0 --repeat-penalty 1.0 (+MTP) - ---------------------------------------…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-05 09:51 · r/LocalLLM
I ran Qwen 3.8 27B on my new MacBook Pro M5 Max and on my RTX 5090 workstation. The Mac held up way better than I expected. - 2026-09-03 23:08 · r/LocalLLM
i made a lot of unofficial tests for different 3 and 4 bit quants of qwen 3.8-27b on my local work on rtx 3090 ti with 96gb ram, and ThinkingCap-Qwen3.6-27B is way better and faster than qwen 3.8-27b, and glm 5.3 and muse spark 1.2, so for me ai benchmarks are useless