Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
TL;DR In this blog post, we present our work on Jagged Flash Attention (JFA) — the attention kernel behind Meta's Generative Ads Model (GEM) — on NVIDIA Blackwell (B200), built...
Read the full story at PyTorch Blog ↗
Timeline · 1 report
- 2026-10-01 22:26 · PyTorch Blog
Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell