Home Projects llama.cpp
llama.cpp
C++

llama.cpp

LLM inference in C/C++

by ggml-org · GitHub
Top 1% most starred in the catalogue
Stars
Forks
License
Created
Last commit
Category
Language
Likely
Self-hostable
ggmlMITC/C++Self-hostable
View on GitHub
In plain words

llama.cpp is the high-performance C/C++ inference engine that underpins most local LLM tools, supporting GGUF models with aggressive quantization across CPUs and GPUs.

From the README

Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
llama.cpp — GitHub preview card
📈 Star history
126k120k
2026-07-122026-08-31
📈 Track llama.cpp

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

LLM inference in C/C++

llama.cpp has 126k stars on GitHub. It has been forked 22.5k times. llama.cpp is written mainly in C/C++. It has been in active development since 2023. llama.cpp is available under the MIT license. Its main topics are ggml.

Read the full guide
Frequently asked questions

What is llama.cpp?

LLM inference in C/C++

Is llama.cpp open source?

llama.cpp is an open-source project. It is released under the MIT license.

Is llama.cpp free?

Yes. llama.cpp is free and open source — you can use, modify and self-host it.

What license does llama.cpp use?

llama.cpp is available under the MIT license.

What language is llama.cpp written in?

llama.cpp is written mainly in C/C++.

🏅 Maintainer of this project?
olud.ai badge — llama.cpp

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=ggml-org-llama-cpp)](https://olud.ai/project/ggml-org-llama-cpp.html)
More badge options →
🧬 Shares DNA with🧬 View the DNA map →

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.