Halving the precision didn't double the speed, because the bottleneck was never the math
https://sathishkumars637097.substack.com/p/fa4-mxfp8-tmem-allocation-problem?r=1jbi98&utm_campaign=post-expanded-share&utm_medium=web
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-21 20:49 · r/learnmachinelearning
Halving the precision didn't double the speed, because the bottleneck was never the math