CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp
MoE speedup, but only for some MoE architectures (like Qwen 35B A3B)
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-10-03 06:18 · r/LocalLLaMA
qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp - 2026-10-02 18:47 · r/LocalLLaMA
CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp