Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. On our internal shapes, FA4 MX8 reaches 2.54 PF/s...
Read the full story at PyTorch Blog ↗
Timeline · 1 report
- 2026-09-16 18:55 · PyTorch Blog
Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell