vllm

47 projets partagent ce topic GitHub

vllm — FunASR ★19.5kvllmllama-cookbook — ★18.5kHalfrost-Field — ★13.2kAI-Research-SKILLs — ★11.2kLMCache — ★10.9kOpenRLHF — ★9.9kinference — ★9.5kMooncake — ★6.1kkserve — ★5.7kUltraRAG — ★5.7kAwesome-LLM-Inference — ★5.4kgpustack — ★5.4ksparrow — ★5.2kllama-swap — ★5.2ksemantic-router — ★5.1ktiny-llm — ★4.4kFastDeploy — ★3.7kcascadeflow — ★3.6kramalama — ★3kvllm-ascend — ★2.5kauto-round — ★1.5kvllm-mlx — ★1.5kInferenceX — ★1.3kBricksLLM — ★1.2kGPTQModel — ★1.2kprometheus-eval — ★1.1ktiny-vllm — ★963UniRL — ★860llmcord — ★819BambooAI — ★784LightCompress — ★736verl-omni — ★676vidur — ★647efficientsam3 — ★644RoboBrain — ★559taOS — ★464nodetool — ★437DeepSeek-OCR-WebUI — ★437smg — ★421topicGPT — ★412super-json-mode — ★396llama-cookbook★ 18.5kHalfrost-Field★ 13.2kAI-Research-SKILLs★ 11.2kLMCache★ 10.9kOpenRLHF★ 9.9kinference★ 9.5kMooncake★ 6.1kkserve★ 5.7kUltraRAG★ 5.7kAwesome-LLM-Inference★ 5.4kgpustack★ 5.4ksparrow★ 5.2kllama-swap★ 5.2ksemantic-router★ 5.1ktiny-llm★ 4.4kFastDeploy★ 3.7kcascadeflow★ 3.6kramalama★ 3kvllm-ascend★ 2.5kauto-round★ 1.5kvllm-mlx★ 1.5kInferenceX★ 1.3kBricksLLM★ 1.2kGPTQModel★ 1.2kprometheus-eval★ 1.1ktiny-vllm★ 963UniRL★ 860llmcord★ 819BambooAI★ 784LightCompress★ 736verl-omni★ 676vidur★ 647efficientsam3★ 644RoboBrain★ 559taOS★ 464nodetool★ 437DeepSeek-OCR-WebUI★ 437smg★ 421topicGPT★ 412super-json-mode★ 396

Les traits relient les membres réellement apparentés entre eux. La taille des points suit les étoiles.

🧬 Membres
FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker…
★ 19.5k
llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with…
★ 18.5k
Halfrost-Field
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field…
★ 13.2k
AI-Research-SKILLs
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills…
★ 11.2k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 10.9k
OpenRLHF
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & …
★ 9.9k
inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and…
★ 9.5k
Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.1k
kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework…
★ 5.7k
UltraRAG
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
★ 5.7k
Awesome-LLM-Inference
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4,…
★ 5.4k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.4k
sparrow
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
★ 5.2k
llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
★ 5.2k
semantic-router
Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
★ 5.1k
tiny-llm
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
★ 4.4k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
cascadeflow
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
★ 3.6k
ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and…
★ 3k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.5k
auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA,…
★ 1.5k
vllm-mlx
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama,…
★ 1.5k
InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 -…
★ 1.3k
BricksLLM
🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get…
★ 1.2k
GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and…
★ 1.2k
prometheus-eval
Evaluate your LLM's response with Prometheus and GPT4 💯
★ 1.1k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 963
UniRL
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
★ 860
llmcord
Make Discord your LLM frontend - Supports any OpenAI compatible API (OpenRouter, Ollama and more)
★ 819
BambooAI
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
★ 784
LightCompress
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video…
★ 736
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 676
vidur
Accurate, large-scale, and extensible simulator for LLM inference Systems
★ 647
efficientsam3
EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation…
★ 644
RoboBrain
[CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official…
★ 559
taOS
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default,…
★ 464
nodetool
The open creative AI workspace
★ 437
DeepSeek-OCR-WebUI
🎨 Ready-to-use DeepSeek-OCR Web UI | Modern Interface | 7 Recognition Modes | Batch Processing | …
★ 437
smg
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM,…
★ 421
topicGPT
TopicGPT: A Prompt-Based Framework for Topic Modeling [NAACL'24]
★ 412
super-json-mode
Low latency JSON generation using LLMs ⚡️
★ 396
TinyLLM
Setup and run a local LLM and Chatbot using consumer grade hardware.
★ 345
Flash-RL
Implementation for FP8/INT8 Rollout for RL training without performence drop.
★ 307
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 128
kube-llmops
★ 122
heron
Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude,…
★ 66 · GitHub ↗
sbproxy
Self-hosted AI gateway and LLM proxy. OpenAI-compatible API for OpenAI, Anthropic, Gemini, Bedrock and 60+…
★ 47 · GitHub ↗
🔗 Familles voisines

Mesuré à partir des topics GitHub communs aux deux projets, pondérés par leur rareté.