ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp
Another day, another Qwen 3.x speedup (prompt processing this time). Soon your Qwen will read your entire project before you can blink! test master t/s PR t/s change pp512 3075.24 3243.60 +5.5% pp4096 3059.87 3221.08 +5.3% tg128 47.62 47.67 +0.1%
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-08 09:04 · r/LocalLLaMA
ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp