AINewsnow

MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash

Fine, pp is still slower, but I'm slowly coming around to the idea of MTP finally being useful on Apple Silicon, and this is the first time I'm seeing a model outperform ds4 (and that's with IngeniousIdiocy's M3U tuning). MTP seems to have no advantage there as was always the case with llama.cpp, u…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-08 15:56 · r/LocalLLaMA
    MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash

More stories

  1. How comparable is a MacBook Pro M5 Pro 64GB 18/20 vs RTX4090 | 128 GB DDR5 — r/LocalLLM
  2. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  4. RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to — r/LocalLLM
  5. Best current R9700 inference engine? — r/LocalLLaMA
  6. Meta's Llama 3.3 70B Now Fits on a Single 48 GB GPU — AlphaSignal
  7. struggling with llama.cpp 2 x dgx spark mtp files start command (unsloth) — r/LocalLLaMA
  8. llama.cpp on the stage — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →