Format-Aware Fusion for Fast FP4 Pretraining
arXiv:2610.00053v1 Announce Type: new Abstract: Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can erase the gain. We present \emph{format-aware fusion}, which co-designs each quantiz…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.LG
Format-Aware Fusion for Fast FP4 Pretraining