Aphrodite Engine is a high-throughput inference server based on vLLM, optimized for serving many users at once with broad quantization and sampling support.
| Category | Inference server |
| Type | Inference server |
| License | AGPL-3.0 |
| Runs locally | Self-hosted |
| Built with | Python |
| Skill level | Advanced |
| Best for | serving many users at high throughput |
Other open-source inference server tools worth comparing:
vLLMHigh-throughput serving for production
TGIHugging Face's production text server
SGLangFast serving with structured outputs
LMDeployToolkit for compressing and serving LLMs
TensorRT-LLMPeak throughput on NVIDIA GPUs
OpenLLMServe any open model as an OpenAI API in one command
KTransformersRun huge MoE models on one consumer GPU
Ray ServeScale model serving across a cluster
BentoMLPackage any model into a production APIAphrodite Engine is free and open-source (AGPL-3.0 license), so you can use, self-host and modify it at no cost.
Yes. Aphrodite Engine is designed to run on your own machine or server, keeping your data private.
Popular open-source alternatives include vLLM, TGI, SGLang. See the comparisons above to choose.
Browse the full directory of open-source AI tools, models and projects — updated daily.
Browse all tools →