A high-throughput and memory-efficient inference and serving engine for LLMs
vLLM is a high-throughput inference and serving engine using PagedAttention to maximize GPU utilization, the default choice for serving open models at scale.
uv pip install vllm
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeA high-throughput and memory-efficient inference and serving engine for LLMs
vllm has 90.5k stars on GitHub. It has been forked 21.4k times. vllm is written mainly in Python. It has been in active development since 2023. vllm is available under the Apache-2.0 license. Its main topics are amd, blackwell, cuda, deepseek.
Read the full guideA high-throughput and memory-efficient inference and serving engine for LLMs
vllm is an open-source project. It is released under the Apache-2.0 license.
Yes. vllm is free and open source — you can use, modify and self-host it.
vllm is available under the Apache-2.0 license.
vllm is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/vllm-project-vllm.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.