ggml-cpu: tiled mul_mat for k-quants by jbooth · Pull Request #27851 · ggml-org/llama.cpp
faster CPU prompt processing: "TL;DR: 3-7x faster CPU mul_mat using VNNI with IMO minimal complexity"
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-26 06:18 · r/LocalLLaMA
ggml-cpu: tiled mul_mat for k-quants by jbooth · Pull Request #27851 · ggml-org/llama.cpp