rocm or vulkan? llama or vllm? Qwen3.8 27b on dual R9700
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
I'm currently running a single r9700 on pci3 16x using Vulkan llama.cpp qwen3.8 27b Q5 MTP, 262k context, q8 kv cache, single stream, and average ~20 tg and ~500 pp. I purchased a 3970x Threadripper and a second r9700 so I can get pci4 16x and plan on going to Q8 and 4 streams. Just looking at what…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-05 12:22 · r/LocalLLM
rocm or vulkan? llama or vllm? Qwen3.8 27b on dual R9700