[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Speed up AI model processing without losing accuracy by using efficient attention techniques in your projects.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytime[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
SageAttention has 3.7k stars on GitHub. It has been forked 495 times. SageAttention is written mainly in Cuda. It has been in active development since 2024. SageAttention is available under the Apache-2.0 license. Its main topics are attention, cuda, efficient-attention, inference-acceleration.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
SageAttention is an open-source project. It is released under the Apache-2.0 license.
Yes. SageAttention is free and open source — you can use, modify and self-host it.
SageAttention is available under the Apache-2.0 license.
SageAttention is written mainly in Cuda.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/thu-ml-sageattention.html)