llama.cpp is the high-performance C/C++ inference engine that underpins most local LLM tools, supporting GGUF models with aggressive quantization across CPUs and GPUs.
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeLLM inference in C/C++
llama.cpp has 126k stars on GitHub. It has been forked 22.5k times. llama.cpp is written mainly in C/C++. It has been in active development since 2023. llama.cpp is available under the MIT license. Its main topics are ggml.
Read the full guideLLM inference in C/C++
llama.cpp is an open-source project. It is released under the MIT license.
Yes. llama.cpp is free and open source — you can use, modify and self-host it.
llama.cpp is available under the MIT license.
llama.cpp is written mainly in C/C++.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/ggml-org-llama-cpp.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.