New FP4 attention kernels for B300 achieve up to 1.69x speedup
The kernels target B300 hardware and offer up to 1.69x speedup over FlashAttention 4 (FA4), as detailed in a tweet from haoailab. The results include benchmark comparisons and kernel design notes.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- OpenAI launches Thailand AI startup accelerator with MHESI
- Engineers increasingly say 'I don't know, Claude wrote this'
- Open source caught up because it's open
- Tencent releases AI model it claims outperforms Z.AI, Moonshot
- OpenAI hires Meta executive to lead Southeast Asia, Australia