Chilled AI & Slow Inference
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
I have a dual GPU set up with 24GB VRAM that runs Qwen3.8:27b Q3 at about 40 T/S. I wanted to test the Q6 quant with a large context window for agentic coding. I dusted down my GMKtec K12 Mini PC with an AMD Radeon 780M iGPU that will comfortably run the Q6 quant with 200k context. It's amazing! Qw…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-16 17:41 · r/LocalLLM
Chilled AI & Slow Inference