llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp
Potentially big speedup for MoE models that don’t fully fit in VRAM. Are you GPU Poor? Show your speedups ;) update https://github.com/ggml-org/llama.cpp/pull/30112 MERGED
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-07 18:15 · r/LocalLLaMA
llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp