inference

94 Projekte teilen dieses GitHub-Topic

inference — vllm ★87.6kinferenceyolov5 — ★57.8kwhisper.cpp — ★52.4kDeepSpeed — ★42.8kColossalAI — ★41.4kmediapipe — ★36.4ksglang — ★30.9kfaster-whisper — ★24.6kncnn — ★23.6kml-engineering — ★18.5knano-vllm — ★14.7kTensorRT — ★13.2kamazon-sagemaker-examples — ★11kLMCache — ★10.9kserver — ★10.9kyolov3 — ★10.6kopenvino — ★10.6krunanywhere-sdks — ★10.3kinference — ★9.5koumi — ★9.4kjetson-inference — ★8.9kargmax-oss-swift — ★6.3kadversarial-robustness-toolbox — ★6.1kMooncake — ★6.1kwhichllm — ★6kvllm-omni — ★5.7kgpustack — ★5.4ksuperduper — ★5.3kcube-studio — ★5.1kTNN — ★4.6kCTranslate2 — ★4.6kcsghub — ★4.2kzml — ★3.9kllm-d — ★3.9kFastVideo — ★3.9kFastDeploy — ★3.7kKuiperInfer — ★3.5koptimum — ★3.5kRapid-MLX — ★3.4kopenvino_notebooks — ★3.2kvllm-ascend — ★2.5kyolov5★ 57.8kwhisper.cpp★ 52.4kDeepSpeed★ 42.8kColossalAI★ 41.4kmediapipe★ 36.4ksglang★ 30.9kfaster-whisper★ 24.6kncnn★ 23.6kml-engineering★ 18.5knano-vllm★ 14.7kTensorRT★ 13.2kamazon-sagemaker-example…★ 11kLMCache★ 10.9kserver★ 10.9kyolov3★ 10.6kopenvino★ 10.6krunanywhere-sdks★ 10.3kinference★ 9.5koumi★ 9.4kjetson-inference★ 8.9kargmax-oss-swift★ 6.3kadversarial-robustness-t…★ 6.1kMooncake★ 6.1kwhichllm★ 6kvllm-omni★ 5.7kgpustack★ 5.4ksuperduper★ 5.3kcube-studio★ 5.1kTNN★ 4.6kCTranslate2★ 4.6kcsghub★ 4.2kzml★ 3.9kllm-d★ 3.9kFastVideo★ 3.9kFastDeploy★ 3.7kKuiperInfer★ 3.5koptimum★ 3.5kRapid-MLX★ 3.4kopenvino_notebooks★ 3.2kvllm-ascend★ 2.5k

Linien verbinden Mitglieder, die messbar miteinander verwandt sind. Die Punktgröße spiegelt die Sterne wider.

🧬 Mitglieder
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 87.6k
yolov5
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and…
★ 57.8k
whisper.cpp
Port of OpenAI's Whisper model in C/C++
★ 52.4k
DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy,…
★ 42.8k
ColossalAI
Making large AI models cheaper, faster and more accessible
★ 41.4k
mediapipe
Cross-platform, customizable ML solutions for live and streaming media.
★ 36.4k
sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 30.9k
faster-whisper
Faster Whisper transcription with CTranslate2
★ 24.6k
ncnn
ncnn is a high-performance neural network inference framework optimized for the mobile platform
★ 23.6k
ml-engineering
Machine Learning Engineering Open Book
★ 18.5k
nano-vllm
Nano vLLM
★ 14.7k
TensorRT
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository…
★ 13.2k
amazon-sagemaker-examples
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using…
★ 11k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 10.9k
server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
★ 10.9k
yolov3
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training,…
★ 10.6k
openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
★ 10.6k
runanywhere-sdks
Production ready toolkit to run AI locally
★ 10.3k
inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and…
★ 9.5k
oumi
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
★ 9.4k
jetson-inference
Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT…
★ 8.9k
argmax-oss-swift
On-device Speech AI for Apple Silicon
★ 6.3k
adversarial-robustness-toolbox
Adversarial Robustness Toolbox (ART) - Python Library for Machine Learning Security - Evasion, Poisoning,…
★ 6.1k
Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6.1k
whichllm
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware…
★ 6k
vllm-omni
A framework for efficient model inference with omni-modality models
★ 5.7k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.4k
superduper
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5.3k
cube-studio
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,…
★ 5.1k
TNN
TNN: developed by Tencent Youtu Lab and Guangying Lab, a uniform deep learning inference framework for…
★ 4.6k
CTranslate2
Fast inference engine for Transformer models
★ 4.6k
csghub
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both…
★ 4.2k
zml
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
★ 3.9k
llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
★ 3.9k
FastVideo
A unified inference and post-training framework for accelerated video generation.
★ 3.9k
FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3.7k
KuiperInfer
★ 3.5k
optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with…
★ 3.5k
Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling.…
★ 3.4k
openvino_notebooks
📚 Jupyter notebook tutorials for OpenVINO™
★ 3.2k
vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
★ 2.5k
huggingface.js
Use Hugging Face with JavaScript
★ 2.5k
ort
Fast ML inference & training for ONNX models in Rust
★ 2.4k
inference
Turn any computer or edge device into a command center for your computer vision projects.
★ 2.4k
cube-studio
★ 2.4k
AI-Engineering.academy
Mastering Applied AI, One Concept at a Time
★ 2.4k
sie
Open-source inference server and production cluster for all the models your agent needs.
★ 2.4k
dstack
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and…
★ 2.2k
any-llm
Communicate with an LLM provider using a single interface
★ 2.1k
DeepSpeed-MII
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
★ 2.1k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
ai-gateway
Manages Unified Access to Generative AI Services built on Envoy Gateway
★ 1.9k
tensorflow_template_application
TensorFlow template application for deep learning
★ 1.9k
agibot_x1_infer
The inference module for AgiBot X1.
★ 1.8k
picolm
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
★ 1.7k
transformer-deploy
Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models…
★ 1.7k
uzu
A high-performance inference engine for AI models
★ 1.7k
nvidia_gpu_exporter
Nvidia GPU exporter for prometheus using nvidia-smi binary
★ 1.5k
xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.…
★ 1.5k
llmgateway
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
★ 1.5k
vllm-mlx
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama,…
★ 1.5k
EmbedAnything
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust…
★ 1.3k
rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
★ 1.3k
kvpress
LLM KV cache compression made easy
★ 1.2k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 963
ims
📚 Introduction to Modern Statistics - A college-level open-source textbook with a modern approach…
★ 938
mlxstudio
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
★ 920
bark.cpp
Suno AI's Bark model in C/C++ for fast text-to-speech generation
★ 866
pipeless
An open-source computer vision framework to build and deploy apps in minutes
★ 851
tensorrt-cpp-api
TensorRT C++ API Tutorial
★ 808
APT
AI Productivity Tool - Free and open source, improve user productivity, and protect privacy and data…
★ 771
GenossGPT
One API for all LLMs either Private or Public (Anthropic, Llama V2, GPT 3.5/4, Vertex, GPT4ALL, HuggingFace…
★ 756
onepanel
The open source, end-to-end computer vision platform. Label, build, train, tune, deploy and automate in a…
★ 730
vidur
Accurate, large-scale, and extensible simulator for LLM inference Systems
★ 647
Deepdive-llama3-from-scratch
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement…
★ 632
distill-sd
Segmind Distilled diffusion
★ 618
YOLO-Patch-Based-Inference
Python library for YOLO small object detection and instance segmentation
★ 554
pinferencia
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
★ 543
aikit
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
★ 533
geti_v2
⚠️ Legacy repository for Geti v2.x. For Geti v3.0+, visit https://github.com/open-edge-platform/geti
★ 484
isaac_ros_pose_estimation
Deep learned, NVIDIA-accelerated 3D object pose estimation
★ 480
SwiftInfer
Efficient AI Inference & Serving
★ 478
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 451
nanoowl
A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.
★ 442
ParaAttention
https://wavespeed.ai/ Context parallel attention that accelerates DiT model inference with dynamic caching
★ 427
KIVI
[ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
★ 423
super-rag
Super performant RAG pipelines for AI apps. Summarization, Retrieve/Rerank and Code Interpreters in one…
★ 395
mlx-serve
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX…
★ 380
gpmp2
Gaussian Process Motion Planner 2
★ 357
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 330
dot-loom
Provider-pluggable orchestration runtime for multi-model AI inference. ( Sakana Fugu style )
★ 257
mstar
A high-performance, universal serving framework for any-to-any models.
★ 60 · GitHub ↗
FerryAI
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process…
★ 56 · GitHub ↗
dgx-spark-inference-stack
Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your…
★ 50 · GitHub ↗
🔗 Verwandte Familien

Gemessen anhand der von beiden Projekten geteilten GitHub-Themen, gewichtet nach der Seltenheit jedes Themas.