AINewsnow

feat: add GLM5Next MTP, optimize by pwilkin · Pull Request #29928 · ggml-org/llama.cpp

now you can use GLM 5 Flash MTP locally

Read the full story at r/LocalLLaMA ↗

Timeline · 3 reports

  1. 2026-10-08 09:04 · r/LocalLLaMA
    ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp
  2. 2026-10-07 18:15 · r/LocalLLaMA
    llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp
  3. 2026-10-07 12:36 · r/LocalLLaMA
    feat: add GLM5Next MTP, optimize by pwilkin · Pull Request #29928 · ggml-org/llama.cpp

More stories

  1. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  2. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  3. Best current R9700 inference engine? — r/LocalLLaMA
  4. Looking for developer-friendly inference providers who give you enough API credits to experiment [D] — r/MachineLearning
  5. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  6. Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M) — r/LocalLLaMA
  7. Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B? — r/LocalLLaMA
  8. China’s open-weight AI models are winning global users. Who is capturing the value? — South China Morning Post Tech

Get the daily brief of stories like this at 6:30 every morning →