WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput
I've released WHIRL (Windows HIP Inference for RDNA LLMs), an open-source (Apache-2.0) LLM inference engine for the AMD Radeon AI PRO R9700 that runs natively on Windows. It's pure C++ and HIP. There's no WSL, no Linux VM and no llama.cpp underneath. You need only the AMD Adrenalin driver to run it…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-03 04:21 · r/LocalLLM
WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput
More stories
- Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
- add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
- Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
- Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU. — r/LocalLLaMA
- FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users. — r/LocalLLaMA
- Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next. — r/LocalLLaMA
- New in llama.cpp: Decision Models — r/LocalLLaMA
- Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience — Latent Space
Get the daily brief of stories like this at 6:30 every morning →