thu-ml/SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
⭐ 3.7k
⑂ 485
Cuda
Apache-2.0
· 2026-01-18推送
3.7k
Watchers
0
贡献者
0
Commits
0
Releases
205
Open Issues
2026-01-18
最近推送
原文
中文
暂无 README