Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp
now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B? (merged after 17h of development) quants: https://huggingface.co/ggml-org/Qwen3.8-Flash-Next-GGUF link to the previous discussion (I deleted the old post to avoid duplicates): https://www.reddit.com/r/LocalLLaMA/comments/1w…
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-10-03 06:18 · r/LocalLLaMA
qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp - 2026-10-02 18:47 · r/LocalLLaMA
CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp - 2026-10-02 10:29 · r/LocalLLaMA
llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp - 2026-10-01 11:18 · r/LocalLLaMA
Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp