llama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
20B-A1B model is coming, good for low VRAM people? https://huggingface.co/deepgrove/maple-preview
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-16 18:36 · r/LocalLLaMA
Enable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp - 2026-09-16 15:51 · r/LocalLLaMA
Release b11003 · ggml-org/llama.cpp - 2026-09-16 09:51 · r/LocalLLaMA
qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp - 2026-09-14 12:11 · r/LocalLLaMA
llama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp