OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
arXiv:2610.09346v1 Announce Type: new Abstract: Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas t…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv cs.CL
OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models