AWQ Looked 10 Points Better on GSM8K Until I Stopped Truncating the Answers
This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.
FlipGate: a release gate that counts per-item answer flips against a measured noise floor, what it found on Qwen2.5-3B, and the bug in my own baseline that I had to fix first. The Problem When you quantise an LLM from bf16 to INT4, the standard metric is accuracy: does the quantised model get the s…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-05 14:30 · DEV Community — Machine Learning
AWQ Looked 10 Points Better on GSM8K Until I Stopped Truncating the Answers