model-serving

16 projects share this GitHub topic

model-serving — vllm ★90.6kmodel-servingBentoML — ★8.8kOlares — ★5.2kAI-Infra-from-Zero-to-Hero — ★4.3kLightLLM — ★4.3kFedML — ★4.1klorax — ★3.8kchitu — ★3kvllm-ascend — ★2.7kopenlake — ★2.4kaici — ★2.1kServerlessLLM — ★711JetStream — ★456predikit — ★400infercrane — ★51dgx-spark-inference-stack — ★51BentoML★ 8.8kOlares★ 5.2kAI-Infra-from-Zero-to-He…★ 4.3kLightLLM★ 4.3kFedML★ 4.1klorax★ 3.8kchitu★ 3kvllm-ascend★ 2.7kopenlake★ 2.4kaici★ 2.1kServerlessLLM★ 711JetStream★ 456predikit★ 400infercrane★ 51dgx-spark-inference-stac…★ 51

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 90.6k
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model…
★ 8.8k
Olares
Open-Source Personal Cloud OS for Always-On Agents
★ 5.2k
AI-Infra-from-Zero-to-Hero
🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for…
★ 4.3k
LightLLM
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its…
★ 4.3k
FedML
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and…
★ 4.1k
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
★ 3.8k
chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and…
★ 3k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.7k
openlake
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
★ 2.4k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
ServerlessLLM
Serverless LLM Serving for Everyone.
★ 711
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 456
predikit
The missing bridge between your ML models and your AI agents.
★ 400
infercrane
Open-source infrastructure for the full inference lifecycle: deploy, observe, scale, optimize, and safely…
★ 51
dgx-spark-inference-stack
Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your…
★ 51
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.