AINewsnow

WHIRL v0.1.3 — native Windows LLM engine for the Radeon AI PRO R9700: up to 2.8× llama.cpp on the same GGUF, same answers bit-for-bit

WHIRL is an open-source (Apache-2.0) inference engine for the AMD Radeon AI PRO R9700 (RDNA 4, 32 GB) on Windows : pure C++/HIP, every kernel included, no WSL or Docker. Just the AMD driver. vs llama.cpp b11214 — same GGUF, same prompts, same R9700, Swift-1.5 27B MXFP4: Decode on coding prompts: 11…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-05 13:29 · r/LocalLLM
    WHIRL v0.1.3 — native Windows LLM engine for the Radeon AI PRO R9700: up to 2.8× llama.cpp on the same GGUF, same answers bit-for-bit

More stories

  1. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  2. LLM Inference Dashboard — r/LocalLLaMA
  3. Photon Announces $4.5M Seed Round to Help Developers Build AI Agents for iMessage and WhatsApp — AI Insider
  4. CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp — r/LocalLLaMA
  6. RTX 5090 local AI setup _ what actually worked for me — r/LocalLLM
  7. GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M — r/LocalLLaMA
  8. Strata looping badly with iq2_xxs — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →