Enable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
one more MTP speedup
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-16 18:36 · r/LocalLLaMA
Enable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp