model-serving

26 projets partagent ce topic GitHub

model-serving — vllm ★87.6kmodel-servingBentoML — ★8.7kkserve — ★5.7kvllm-omni — ★5.7kOlares — ★5.1kDeep-Learning-in-Production — ★4.4kAI-Infra-from-Zero-to-Hero — ★4.2kLightLLM — ★4.2kFedML — ★4.1klorax — ★3.8kchitu — ★3.1kvllm-ascend — ★2.5kopenlake — ★2.3kenvd — ★2.2kaici — ★2.1kmlrun — ★1.7krtp-llm — ★1.3ktruss — ★1.2kZhiLight — ★907ServerlessLLM — ★695pinferencia — ★543JetStream — ★451predikit — ★420BentoDiffusion — ★388swiftLLM — ★330dgx-spark-inference-stack — ★50BentoML★ 8.7kkserve★ 5.7kvllm-omni★ 5.7kOlares★ 5.1kDeep-Learning-in-Product…★ 4.4kAI-Infra-from-Zero-to-He…★ 4.2kLightLLM★ 4.2kFedML★ 4.1klorax★ 3.8kchitu★ 3.1kvllm-ascend★ 2.5kopenlake★ 2.3kenvd★ 2.2kaici★ 2.1kmlrun★ 1.7krtp-llm★ 1.3ktruss★ 1.2kZhiLight★ 907ServerlessLLM★ 695pinferencia★ 543JetStream★ 451predikit★ 420BentoDiffusion★ 388swiftLLM★ 330dgx-spark-inference-stac…★ 50 · GitHub ↗

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 87.6k
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model…
★ 8.7k
kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework…
★ 5.7k
vllm-omni
A framework for efficient model inference with omni-modality models
★ 5.7k
Olares
Open-Source Personal Cloud OS for Always-On Agents
★ 5.1k
Deep-Learning-in-Production
In this repository, I will share some useful notes and references about deploying deep learning-based models…
★ 4.4k
AI-Infra-from-Zero-to-Hero
🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for…
★ 4.2k
LightLLM
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its…
★ 4.2k
FedML
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and…
★ 4.1k
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
★ 3.8k
chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and…
★ 3.1k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.5k
openlake
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
★ 2.3k
envd
🏕️ Reproducible development environment for humans and agents
★ 2.2k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
mlrun
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across…
★ 1.7k
rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
★ 1.3k
truss
The simplest way to serve AI/ML models in production
★ 1.2k
ZhiLight
A highly optimized LLM inference acceleration engine for Llama and its variants.
★ 907
ServerlessLLM
Serverless LLM Serving for Everyone.
★ 695
pinferencia
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
★ 543
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 451
predikit
The missing bridge between your ML models and your AI agents.
★ 420
BentoDiffusion
BentoDiffusion: A collection of diffusion models served with BentoML
★ 388
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 330
dgx-spark-inference-stack
Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your…
★ 50 · GitHub ↗
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.