One vLLM flag changed my AWQ-vs-fp16 cost result by 73 percentage points
This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.
I wanted to know whether 4-bit quantization actually saves money, not memory. So I measured Qwen2.5-1.5B against its AWQ version on a T4, in dollars per million output tokens. First run said quantization was MORE expensive: batch 1 AWQ +24.8% vs fp16 batch 8 AWQ +14.0% batch 32 AWQ +16.0% batch 128…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-21 09:06 · r/LocalLLM
One vLLM flag changed my AWQ-vs-fp16 cost result by 73 percentage points