llm-serving

17 proyectos comparten este topic de GitHub

llm-serving — vllm ★90.6kllm-servingray — ★43.7kllm-action — ★25kTensorRT-LLM — ★14.5kOpenLLM — ★12.5kBentoML — ★8.8kgpustack — ★5.6klorax — ★3.8kFastDeploy — ★3.7kchitu — ★3kvllm-ascend — ★2.7kMoBA — ★2.2kaici — ★2.1kparallax — ★1.4khelix — ★793LLM-FineTuning-Large-Language-Models — ★578sndr_core_engine — ★131ray★ 43.7kllm-action★ 25kTensorRT-LLM★ 14.5kOpenLLM★ 12.5kBentoML★ 8.8kgpustack★ 5.6klorax★ 3.8kFastDeploy★ 3.7kchitu★ 3kvllm-ascend★ 2.7kMoBA★ 2.2kaici★ 2.1kparallax★ 1.4khelix★ 793LLM-FineTuning-Large-Lan…★ 578sndr_core_engine★ 131

Las líneas conectan a los miembros que están mediblemente relacionados entre sí. El tamaño de los puntos refleja las estrellas.

🧬 Miembros
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 90.6k
ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for…
★ 43.7k
llm-action
★ 25k
TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and…
★ 14.5k
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12.5k
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model…
★ 8.8k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.6k
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
★ 3.8k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and…
★ 3k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.7k
MoBA
MoBA: Mixture of Block Attention for Long-Context LLMs
★ 2.2k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
parallax
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere
★ 1.4k
helix
♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude,…
★ 793
LLM-FineTuning-Large-Language-Models
LLM (Large Language Model) FineTuning
★ 578
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 131
🔗 Familias relacionadas

Medido a partir de los temas de GitHub compartidos por ambos proyectos, ponderado por cuán raros son cada uno de los temas.