inference

74 progetti condividono questo topic GitHub

inference — vllm ★90.6kinferencewhisper.cpp — ★51.8kColossalAI — ★41.4ksglang — ★33kfaster-whisper — ★24.3kml-engineering — ★18.9knano-vllm — ★15.2kTensorRT — ★13.2kLMCache — ★11.6kopenvino — ★10.8krunanywhere-sdks — ★10.3kinference — ★9.5koumi — ★9.4kMooncake — ★6.4kargmax-oss-swift — ★6.3kvllm-omni — ★5.6kgpustack — ★5.6ksemantic-router — ★5.5kllm-d — ★4.4kcsghub — ★4.1kzml — ★3.9kFastVideo — ★3.8kFastDeploy — ★3.7kRapid-MLX — ★3.6koptimum — ★3.4kopenvino_notebooks — ★3.2ksie — ★2.9kvllm-ascend — ★2.7khuggingface.js — ★2.5kinference — ★2.4kort — ★2.4kAI-Engineering.academy — ★2.4kany-llm — ★2.2kDeepSpeed-MII — ★2.1kaici — ★2.1kai-gateway — ★2kpicolm — ★1.9ktensorflow_template_application — ★1.9kagibot_x1_infer — ★1.8ktransformer-deploy — ★1.7kuzu — ★1.7kwhisper.cpp★ 51.8kColossalAI★ 41.4ksglang★ 33kfaster-whisper★ 24.3kml-engineering★ 18.9knano-vllm★ 15.2kTensorRT★ 13.2kLMCache★ 11.6kopenvino★ 10.8krunanywhere-sdks★ 10.3kinference★ 9.5koumi★ 9.4kMooncake★ 6.4kargmax-oss-swift★ 6.3kvllm-omni★ 5.6kgpustack★ 5.6ksemantic-router★ 5.5kllm-d★ 4.4kcsghub★ 4.1kzml★ 3.9kFastVideo★ 3.8kFastDeploy★ 3.7kRapid-MLX★ 3.6koptimum★ 3.4kopenvino_notebooks★ 3.2ksie★ 2.9kvllm-ascend★ 2.7khuggingface.js★ 2.5kinference★ 2.4kort★ 2.4kAI-Engineering.academy★ 2.4kany-llm★ 2.2kDeepSpeed-MII★ 2.1kaici★ 2.1kai-gateway★ 2kpicolm★ 1.9ktensorflow_template_appl…★ 1.9kagibot_x1_infer★ 1.8ktransformer-deploy★ 1.7kuzu★ 1.7k

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 90.6k
whisper.cpp
Port of OpenAI's Whisper model in C/C++
★ 51.8k
ColossalAI
Making large AI models cheaper, faster and more accessible
★ 41.4k
sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 33k
faster-whisper
Faster Whisper transcription with CTranslate2
★ 24.3k
ml-engineering
Machine Learning Engineering Open Book
★ 18.9k
nano-vllm
Nano vLLM
★ 15.2k
TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository…
★ 13.2k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 11.6k
openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
★ 10.8k
runanywhere-sdks
Production ready toolkit to run AI locally
★ 10.3k
inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and…
★ 9.5k
oumi
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
★ 9.4k
Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.4k
argmax-oss-swift
On-device Speech AI for Apple Silicon
★ 6.3k
vllm-omni
A framework for efficient model inference with omni-modality models
★ 5.6k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.6k
semantic-router
A programmable Mixture-of-Models router for heterogeneous LLM inference
★ 5.5k
llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
★ 4.4k
csghub
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both…
★ 4.1k
zml
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
★ 3.9k
FastVideo
A unified inference and post-training framework for accelerated video generation.
★ 3.8k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling.…
★ 3.6k
optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with…
★ 3.4k
openvino_notebooks
📚 Jupyter notebook tutorials for OpenVINO™
★ 3.2k
sie
Open-source inference server and production cluster for all the models your agent needs.
★ 2.9k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.7k
huggingface.js
Use Hugging Face with JavaScript
★ 2.5k
inference
Turn any computer or edge device into a command center for your computer vision projects.
★ 2.4k
ort
Fast ML inference & training for ONNX models in Rust
★ 2.4k
AI-Engineering.academy
Mastering Applied AI, One Concept at a Time
★ 2.4k
any-llm
Communicate with an LLM provider using a single interface
★ 2.2k
DeepSpeed-MII
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
★ 2.1k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
ai-gateway
Manages Unified Access to Generative AI Services built on Envoy Gateway
★ 2k
picolm
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
★ 1.9k
tensorflow_template_application
TensorFlow template application for deep learning
★ 1.9k
agibot_x1_infer
The inference module for AgiBot X1.
★ 1.8k
transformer-deploy
Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models…
★ 1.7k
uzu
A high-performance inference engine for AI models
★ 1.7k
llmgateway
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
★ 1.6k
nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML
★ 1.5k
xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.…
★ 1.5k
EmbedAnything
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust…
★ 1.3k
rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
★ 1.3k
kvpress
LLM KV cache compression made easy
★ 1.2k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1.1k
mlx-serve
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX…
★ 1k
mlxstudio
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
★ 960
ims
📚 Introduction to Modern Statistics - A college-level open-source textbook with a modern approach…
★ 938
bark.cpp
Suno AI's Bark model in C/C++ for fast text-to-speech generation
★ 866
tensorrt-cpp-api
TensorRT C++ API Tutorial
★ 807
APT
AI Productivity Tool - Free and open source, improve user productivity, and protect privacy and data…
★ 769
vidur
Accurate, large-scale, and extensible simulator for LLM inference Systems
★ 642
Deepdive-llama3-from-scratch
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement…
★ 631
distill-sd
Segmind Distilled diffusion
★ 618
optimum-intel
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
★ 606
YOLO-Patch-Based-Inference
Python library for YOLO small object detection and instance segmentation
★ 553
aikit
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
★ 537
SwiftInfer
Efficient AI Inference & Serving
★ 478
isaac_ros_pose_estimation
Deep learned, NVIDIA-accelerated 3D object pose estimation
★ 475
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 456
nanoowl
A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.
★ 440
KIVI
[ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
★ 431
super-rag
Super performant RAG pipelines for AI apps. Summarization, Retrieve/Rerank and Code Interpreters in one…
★ 395
TensorSharp
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based…
★ 388
gpmp2
Gaussian Process Motion Planner 2
★ 357
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 329
dot-loom
Provider-pluggable orchestration runtime for multi-model AI inference. ( Sakana Fugu style )
★ 197
mstar
A high-performance, universal serving framework for any-to-any models.
★ 82
infercrane
Open-source infrastructure for the full inference lifecycle: deploy, observe, scale, optimize, and safely…
★ 51
dgx-spark-inference-stack
Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your…
★ 51
FerryAI
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process…
★ 41
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.