AINewsnow

vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (… · ggml-org/llama.cpp@70c4e15

This has made a massive improvement in performance on my 7900XTX before: | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: | | gemma4 26B.A4B…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-26 06:18 · r/LocalLLaMA
    ggml-cpu: tiled mul_mat for k-quants by jbooth · Pull Request #27851 · ggml-org/llama.cpp
  2. 2026-09-24 13:37 · r/LocalLLaMA
    vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (… · ggml-org/llama.cpp@70c4e15

More stories

  1. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  2. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  3. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  4. Targeting Applied AI / ML Engineer roles. Need ruthless feedback on my architecture and metrics. — r/deeplearning
  5. KV Cache Math: Why Llama 3.1 8B at 128K Context Won't Fit in 24GB — DEV Community — Machine Learning
  6. 42x Faster Prompt Lookup Drafting in llama.cpp — r/LocalLLaMA
  7. Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM
  8. Qwen3.8 27B on Intel X7 358h + B390, with pi + llama.cpp surprised by its own RAM speed — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →