Home Projects tiny-vllm
tiny-vllm
C++

tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

by jmaczan · GitHub
Stars
Forks
License
Created
Last commit
Category
Language
aiattentionbatchingApache-2.0C++
View on GitHub
In plain words

Learn to build a fast engine for running large language models with C++ and CUDA.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
tiny-vllm — GitHub preview card
📈 Star history
1.1k0.9k
2026-07-202026-08-31
📈 Track tiny-vllm

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

tiny-vllm has 1.1k stars on GitHub. It has been forked 84 times. tiny-vllm is written mainly in C++. It has been in active development since 2026. tiny-vllm is available under the Apache-2.0 license. Its main topics are ai, attention, batching, course.

Frequently asked questions

What is tiny-vllm?

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

Is tiny-vllm open source?

tiny-vllm is an open-source project. It is released under the Apache-2.0 license.

Is tiny-vllm free?

Yes. tiny-vllm is free and open source — you can use, modify and self-host it.

What license does tiny-vllm use?

tiny-vllm is available under the Apache-2.0 license.

What language is tiny-vllm written in?

tiny-vllm is written mainly in C++.

🏅 Maintainer of this project?
olud.ai badge — tiny-vllm

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=jmaczan-tiny-vllm)](https://olud.ai/project/jmaczan-tiny-vllm.html)
More badge options →
🧬 Related projects🧬 View the DNA map →