Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Yet another vLLM fork thread here, but this time its for older INT8-centric hardware. This is a complete INT8 serving stack for Qwen3.8 27B based on vLLM, AITER, and a 27B GPTQ INT8 quant w/ DFlash2 . Its not just another vibed autoresearch loop. No, vLLM ships with very little int8 support, and th…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-26 20:54 · r/LocalLLaMA
Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork