thu-ml

thu-ml/SageAttention

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

⭐ 3.7k ⑂ 485 Cuda Apache-2.0 · 2026-01-18推送
3.7k
Watchers
0
贡献者
0
Commits
0
Releases
205
Open Issues
2026-01-18
最近推送
原文 中文
暂无 README