AINewsnow

ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp

Another day, another Qwen 3.x speedup (prompt processing this time). Soon your Qwen will read your entire project before you can blink! test master t/s PR t/s change pp512 3075.24 3243.60 +5.5% pp4096 3059.87 3221.08 +5.3% tg128 47.62 47.67 +0.1%

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-08 09:04 · r/LocalLLaMA
    ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp

More stories

  1. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  2. Success with Qwen3.8 27B GSQ-RCO-IQ3_S on 16GB VRAM — r/LocalLLM
  3. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  4. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  5. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  6. Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M) — r/LocalLLaMA
  7. Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B? — r/LocalLLaMA
  8. Best uncensored version of Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S ? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →