The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
BentoML packages models, code and dependencies into a reproducible artifact and serves it as a scalable API, with adaptive batching built in.
pip install torch transformers # additional dependencies for local run
Documentation ↗Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeThe easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
BentoML has 8.8k stars on GitHub. It has been forked 1k times. BentoML is written mainly in Python. It has been in active development since 2019. BentoML is available under the Apache-2.0 license. Its main topics are ai-inference, deep-learning, generative-ai, inference-platform.
Read the full guideThe easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
BentoML is an open-source project. It is released under the Apache-2.0 license.
Yes. BentoML is free and open source — you can use, modify and self-host it.
BentoML is available under the Apache-2.0 license.
BentoML is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/bentoml-bentoml.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.