The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks
TL;DR: Local agent loop, ~21 days, one RTX 3090. Task was pretty much "build a CUDA inference engine for optimized for yourself on this GPU arch." Got working kernels and benches, not a win over llama.cpp. ~12 human messages. Compaction ate ~83 hours. Old joke: you don’t criticize how well the bear…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-20 18:26 · r/LocalLLaMA
The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks