llm-serving

23 progetti condividono questo topic GitHub

llm-serving — vllm ★87.6kllm-servingray — ★43.4kllm-action — ★24.8kTensorRT-LLM — ★14.2kOpenLLM — ★12.4kskypilot — ★10.4kBentoML — ★8.7kgpustack — ★5.4ksuperduper — ★5.3klorax — ★3.8kFastDeploy — ★3.7kchitu — ★3.1kvllm-ascend — ★2.5kMoBA — ★2.2kaici — ★2.1kparallax — ★1.4krtp-llm — ★1.3kZhiLight — ★907helix — ★794LLM-FineTuning-Large-Language-Models — ★576SwiftInfer — ★478swiftLLM — ★330sndr_core_engine — ★128ray★ 43.4kllm-action★ 24.8kTensorRT-LLM★ 14.2kOpenLLM★ 12.4kskypilot★ 10.4kBentoML★ 8.7kgpustack★ 5.4ksuperduper★ 5.3klorax★ 3.8kFastDeploy★ 3.7kchitu★ 3.1kvllm-ascend★ 2.5kMoBA★ 2.2kaici★ 2.1kparallax★ 1.4krtp-llm★ 1.3kZhiLight★ 907helix★ 794LLM-FineTuning-Large-Lan…★ 576SwiftInfer★ 478swiftLLM★ 330sndr_core_engine★ 128

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 87.6k
ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for…
★ 43.4k
llm-action
★ 24.8k
TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and…
★ 14.2k
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12.4k
skypilot
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer,…
★ 10.4k
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model…
★ 8.7k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.4k
superduper
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5.3k
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
★ 3.8k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and…
★ 3.1k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.5k
MoBA
MoBA: Mixture of Block Attention for Long-Context LLMs
★ 2.2k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
parallax
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere
★ 1.4k
rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
★ 1.3k
ZhiLight
A highly optimized LLM inference acceleration engine for Llama and its variants.
★ 907
helix
♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude,…
★ 794
LLM-FineTuning-Large-Language-Models
LLM (Large Language Model) FineTuning
★ 576
SwiftInfer
Efficient AI Inference & Serving
★ 478
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 330
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 128
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.