RDNA4 owners on ComfyUI: I made SageAttention actually fast on the 9070 XT, here's the build plus every caveat
This story is from 2026-10-01. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: A drop-in sageattention build for RX 9070 / 9070 XT on Windows . The common case (fp16, head_dim 128) runs on an fp8 attention kernel I wrote by hand in HIP. Against the existing gfx12 port (SageAttention PR #368), each attention call is 1.07–1.21× faster non-causal and 1.72–1.96× faster cau…
Read the full story at r/StableDiffusion ↗
Timeline · 1 report
- 2026-10-01 23:08 · r/StableDiffusion
RDNA4 owners on ComfyUI: I made SageAttention actually fast on the 9070 XT, here's the build plus every caveat