vllm

35 progetti condividono questo topic GitHub

vllm — FunASR ★20.1kvllmllama-cookbook — ★18.6kHalfrost-Field — ★13.2kLMCache — ★11.6kOpenRLHF — ★10kinference — ★9.5kMooncake — ★6.4kUltraRAG — ★5.7kgpustack — ★5.6kAwesome-LLM-Inference — ★5.5ksemantic-router — ★5.5ksparrow — ★5.2ktiny-llm — ★4.5kcascadeflow — ★4kFastDeploy — ★3.7kramalama — ★3kvllm-ascend — ★2.7kInferenceX — ★1.6kvllm-mlx — ★1.6kBricksLLM — ★1.2ktiny-vllm — ★1.1kverl-omni — ★915UniRL — ★820BambooAI — ★786LightCompress — ★743RoboBrain — ★559taOS — ★503super-json-mode — ★396TinyLLM — ★348Flash-RL — ★306sndr_core_engine — ★131kube-llmops — ★101heron — ★95infercrane — ★51sbproxy — ★51llama-cookbook★ 18.6kHalfrost-Field★ 13.2kLMCache★ 11.6kOpenRLHF★ 10kinference★ 9.5kMooncake★ 6.4kUltraRAG★ 5.7kgpustack★ 5.6kAwesome-LLM-Inference★ 5.5ksemantic-router★ 5.5ksparrow★ 5.2ktiny-llm★ 4.5kcascadeflow★ 4kFastDeploy★ 3.7kramalama★ 3kvllm-ascend★ 2.7kInferenceX★ 1.6kvllm-mlx★ 1.6kBricksLLM★ 1.2ktiny-vllm★ 1.1kverl-omni★ 915UniRL★ 820BambooAI★ 786LightCompress★ 743RoboBrain★ 559taOS★ 503super-json-mode★ 396TinyLLM★ 348Flash-RL★ 306sndr_core_engine★ 131kube-llmops★ 101heron★ 95infercrane★ 51sbproxy★ 51

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker…
★ 20.1k
llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with…
★ 18.6k
Halfrost-Field
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field…
★ 13.2k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 11.6k
OpenRLHF
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & …
★ 10k
inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and…
★ 9.5k
Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.4k
UltraRAG
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
★ 5.7k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.6k
Awesome-LLM-Inference
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4,…
★ 5.5k
semantic-router
A programmable Mixture-of-Models router for heterogeneous LLM inference
★ 5.5k
sparrow
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
★ 5.2k
tiny-llm
learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen
★ 4.5k
cascadeflow
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
★ 4k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and…
★ 3k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.7k
InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 -…
★ 1.6k
vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX,…
★ 1.6k
BricksLLM
🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get…
★ 1.2k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1.1k
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 915
UniRL
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
★ 820
BambooAI
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
★ 786
LightCompress
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video…
★ 743
RoboBrain
[CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official…
★ 559
taOS
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default,…
★ 503
super-json-mode
Low latency JSON generation using LLMs ⚡️
★ 396
TinyLLM
Setup and run a local LLM and Chatbot using consumer grade hardware.
★ 348
Flash-RL
Implementation for FP8/INT8 Rollout for RL training without performence drop.
★ 306
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 131
kube-llmops
★ 101
heron
Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude,…
★ 95
infercrane
Open-source infrastructure for the full inference lifecycle: deploy, observe, scale, optimize, and safely…
★ 51
sbproxy
Open source Enterprise AI Gateway for API, MCP and agent, and AI model traffic. One Apache-2.0 binary: 72…
★ 51
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.