vllm + p2p driver hack + qwen 3.8 27B vs llamacpp + qwen flash next ?
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
Hi everyone I'm running four rtx 4090, 64GB ram, on a threadripper pro motherboard so all PCIe x16 ports, as a homelab machine for coding. I was migrating from vllm + qwen 3.8 27B (fp8+256k kv cache) to llamacpp + qwen flash next iq4xs + 8 bit cache 200k kv cache... until someone had to ruin my mig…
Read the full story at r/LocalLLaMA ↗
Timeline · 3 reports
- 2026-09-08 15:53 · r/LocalLLM
Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s - 2026-09-07 15:25 · r/LocalLLaMA
Are you running Qwen 3.8 27b or Qwen Flash Next? - 2026-09-06 23:00 · r/LocalLLaMA
vllm + p2p driver hack + qwen 3.8 27B vs llamacpp + qwen flash next ?