Home Projects Star-Attention
Star-Attention
Python

Star-Attention

Efficient LLM Inference over Long Sequences

by NVIDIA · GitHub
Stars
Forks
License
Created
Last commit
Language
attention-mechanismlarge-language-modelsllm-inferenceApache-2.0Python
View on GitHub
In plain words

Run long sequence AI models more efficiently with Star Attention, speeding up their response times significantly.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
Star-Attention — GitHub preview card
📈 Star history
392391
2026-07-202026-08-31
📈 Track Star-Attention

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Efficient LLM Inference over Long Sequences

Star-Attention has 392 stars on GitHub. It has been forked 24 times. Star-Attention is written mainly in Python. It has been in active development since 2024. Star-Attention is available under the Apache-2.0 license. Its main topics are attention-mechanism, large-language-models, llm-inference.

Frequently asked questions

What is Star-Attention?

Efficient LLM Inference over Long Sequences

Is Star-Attention open source?

Star-Attention is an open-source project. It is released under the Apache-2.0 license.

Is Star-Attention free?

Yes. Star-Attention is free and open source — you can use, modify and self-host it.

What license does Star-Attention use?

Star-Attention is available under the Apache-2.0 license.

What language is Star-Attention written in?

Star-Attention is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — Star-Attention

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=nvidia-star-attention)](https://olud.ai/project/nvidia-star-attention.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →
Medusa
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
2.8k · llm
sharesllm-inference
MultiModalMamba
A novel implementation of fusing ViT with Mamba into a fast, agile, and high performance Multi-…
473 · ai
sharesattention-mechanism
sign-language-translator
Python library & framework to build custom translators for the hearing-impaired and translate b…
359 · artificial-intelligence
sharesattention-mechanism
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
77.4k · ai-chat
sharesllm-inference
reformer-pytorch
Reformer, the efficient Transformer, in Pytorch
2.2k · artificial-intelligence
sharesattention-mechanism
x-transformers
A concise but complete full-attention transformer with a set of promising experimental features…
5.9k · artificial-intelligence
sharesattention-mechanism
flamingo-pytorch
Implementation of 🦩 Flamingo, state-of-the-art few-shot visual question answering attention net…
1.3k · artificial-intelligence
sharesattention-mechanism
TimeSformer-pytorch
Implementation of TimeSformer from Facebook AI, a pure attention-based solution for video class…
729 · artificial-intelligence
sharesattention-mechanism
nndl
邱锡鹏《神经网络与深度学习》(蒲公英书)理论书 v2 与通识版
19k · attention-mechanism
sharesattention-mechanism
GeoTransformer
[CVPR2022] Geometric Transformer for Fast and Robust Point Cloud Registration
966 · attention-mechanism
sharesattention-mechanism

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.