AINewsnow

40Gb VRAM and 128Gb RAM - Which MoE should I try out?

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

Yes I know Qwen 3.8 27b is probably best choice. And I'm currently using and loving it! But I'm curious to try out bigger models and see how they run. Where would you start? Is llama.cpp best for stuff like moe models or would you use something else? Specs: GPU: 2x 3080 20gb CPU: Xeon 2667 v4 RAM:…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-10 09:13 · r/LocalLLM
    40Gb VRAM and 128Gb RAM - Which MoE should I try out?

More stories

  1. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  2. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  3. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  4. My Version of Jev running locally, playing doom. — r/LocalLLM
  5. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  6. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  7. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  8. Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →