CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp
Another day, another Qwen Flash Next speedup
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-20 16:03 · r/LocalLLaMA
CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp