Home Projects SageAttention
SageAttention
Cuda

SageAttention

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

by thu-ml · GitHub
Stars
Forks
License
Created
Last commit
Category
Cuda
Language
attentioncudaefficient-attentionApache-2.0Cuda
View on GitHub
In plain words

Speed up AI model processing without losing accuracy by using efficient attention techniques in your projects.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
SageAttention — GitHub preview card
📈 Star history
4k3k
2026-07-072026-08-31
📈 Track SageAttention

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

SageAttention has 3.7k stars on GitHub. It has been forked 495 times. SageAttention is written mainly in Cuda. It has been in active development since 2024. SageAttention is available under the Apache-2.0 license. Its main topics are attention, cuda, efficient-attention, inference-acceleration.

Frequently asked questions

What is SageAttention?

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Is SageAttention open source?

SageAttention is an open-source project. It is released under the Apache-2.0 license.

Is SageAttention free?

Yes. SageAttention is free and open source — you can use, modify and self-host it.

What license does SageAttention use?

SageAttention is available under the Apache-2.0 license.

What language is SageAttention written in?

SageAttention is written mainly in Cuda.

🏅 Maintainer of this project?
olud.ai badge — SageAttention

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=thu-ml-sageattention)](https://olud.ai/project/thu-ml-sageattention.html)
More badge options →
🧬 Related projects🧬 View the DNA map →