40Gb VRAM and 128Gb RAM - Which MoE should I try out?
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
Yes I know Qwen 3.8 27b is probably best choice. And I'm currently using and loving it! But I'm curious to try out bigger models and see how they run. Where would you start? Is llama.cpp best for stuff like moe models or would you use something else? Specs: GPU: 2x 3080 20gb CPU: Xeon 2667 v4 RAM:…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-10 09:13 · r/LocalLLM
40Gb VRAM and 128Gb RAM - Which MoE should I try out?