llm-inference

65 projects share this GitHub topic

llm-inference — gpt4all ★77.4kllm-inferenceray — ★43.7kgitleaks — ★29kllm-action — ★25klitgpt — ★13.6kOpenLLM — ★12.5kopenvino — ★10.8kPowerInfer — ★9.8kBentoML — ★8.8klmdeploy — ★8kplano — ★7kkimi-k3-in-c — ★6.9kturbo-fieldfare — ★6.5kflashinfer — ★6.3kcactus — ★6kgpustack — ★5.6klemonade — ★5.6kAwesome-LLM-Inference — ★5.5keko — ★5koptillm — ★4.3kGenerativeAIExamples — ★4.2klorax — ★3.8kAI-Engineer-Headquarters — ★3.7kdistributed-llama — ★3.1kMedusa — ★2.8kEAGLE — ★2.5knanocoder — ★2.4kneuron-ai — ★2.1kaici — ★2.1kllama2-webui — ★1.9kbeta9 — ★1.8kreact-native-executorch — ★1.7kxllm — ★1.5kBrowserAI — ★1.4kawesome-ai-web-search — ★1.4kAtomic-Chat — ★1.4kawesome-hacking-lists — ★1.4kLeanCopilot — ★1.3ktiny-vllm — ★1.1kblast — ★779LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing — ★731ray★ 43.7kgitleaks★ 29kllm-action★ 25klitgpt★ 13.6kOpenLLM★ 12.5kopenvino★ 10.8kPowerInfer★ 9.8kBentoML★ 8.8klmdeploy★ 8kplano★ 7kkimi-k3-in-c★ 6.9kturbo-fieldfare★ 6.5kflashinfer★ 6.3kcactus★ 6kgpustack★ 5.6klemonade★ 5.6kAwesome-LLM-Inference★ 5.5keko★ 5koptillm★ 4.3kGenerativeAIExamples★ 4.2klorax★ 3.8kAI-Engineer-Headquarters★ 3.7kdistributed-llama★ 3.1kMedusa★ 2.8kEAGLE★ 2.5knanocoder★ 2.4kneuron-ai★ 2.1kaici★ 2.1kllama2-webui★ 1.9kbeta9★ 1.8kreact-native-executorch★ 1.7kxllm★ 1.5kBrowserAI★ 1.4kawesome-ai-web-search★ 1.4kAtomic-Chat★ 1.4kawesome-hacking-lists★ 1.4kLeanCopilot★ 1.3ktiny-vllm★ 1.1kblast★ 779LLM-PowerHouse-A-Curated…★ 731

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77.4k
ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for…
★ 43.7k
gitleaks
Find secrets with Gitleaks 🔑
★ 29k
llm-action
★ 25k
litgpt
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
★ 13.6k
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12.5k
openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
★ 10.8k
PowerInfer
High-speed Large Language Model Serving for Local Deployment
★ 9.8k
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model…
★ 8.8k
lmdeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
★ 8k
plano
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent…
★ 7k
kimi-k3-in-c
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS,…
★ 6.9k
turbo-fieldfare
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
★ 6.5k
flashinfer
FlashInfer: Kernel Library for LLM Serving
★ 6.3k
cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
★ 6k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.6k
lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and…
★ 5.6k
Awesome-LLM-Inference
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4,…
★ 5.5k
eko
Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai
★ 5k
optillm
Optimizing inference proxy for LLMs
★ 4.3k
GenerativeAIExamples
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
★ 4.2k
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
★ 3.8k
AI-Engineer-Headquarters
A collection of scientific methods, processes, algorithms, and systems to build stories & models.
★ 3.7k
distributed-llama
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More…
★ 3.1k
Medusa
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8k
EAGLE
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2.5k
nanocoder
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own…
★ 2.4k
neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect…
★ 2.1k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
llama2-webui
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper`…
★ 1.9k
beta9
Ultrafast serverless GPU inference, sandboxes, and background jobs
★ 1.8k
react-native-executorch
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1.7k
xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.…
★ 1.5k
BrowserAI
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser
★ 1.4k
awesome-ai-web-search
List of software that allows searching the web with the assistance of AI:…
★ 1.4k
Atomic-Chat
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your…
★ 1.4k
awesome-hacking-lists
A curated collection of top-tier penetration testing tools and productivity utilities across multiple…
★ 1.4k
LeanCopilot
LLMs as Copilots for Theorem Proving in Lean
★ 1.3k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1.1k
blast
Open-source VMs-as-a-service
★ 779
LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing
LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for…
★ 731
MiniSearch
Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo:…
★ 586
LLM-FineTuning-Large-Language-Models
LLM (Large Language Model) FineTuning
★ 578
krasis
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM…
★ 516
taOS
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default,…
★ 503
incognide
Explore the unknown, build the future, own your data.
★ 468
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 456
sagify
LLMs and Machine Learning done easily
★ 442
Context-Engine
Context-Engine MCP - Agentic Context Compression Suite
★ 400
Star-Attention
Efficient LLM Inference over Long Sequences
★ 392
embedding_studio
Embedding Studio is a framework which allows you transform your Vector Database into a feature-rich Search…
★ 382
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models,…
★ 379
MoE-Infinity
PyTorch library for cost-effective, fast and easy serving of MoE models.
★ 352
syncode
Efficient and general syntactical decoding for Large Language Models
★ 339
monocle
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI…
★ 337
ecologits
🌱 EcoLogits tracks the energy consumption and environmental footprint of using generative AI models…
★ 320
picollm
On-device LLM Inference Powered by X-Bit Quantization
★ 317
leanctx
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code.…
★ 316
LLM-TPU
Run generative AI models in sophgo BM1684X/BM1688
★ 305
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 131
llmtrace
Zero-code LLM security & observability proxy. Real-time prompt injection detection, PII scanning, and cost…
★ 52
infercrane
Open-source infrastructure for the full inference lifecycle: deploy, observe, scale, optimize, and safely…
★ 51
llmff
Bounded, inspectable LLM inference pipelines from declared YAML — runs offline against Ollama or any…
★ 46
halofpx
Run Ornith, Qwen, Nemotron & DeepSeek optimized on AMD Strix Halo — unified OpenAI-compatible server with…
★ 41 · GitHub ↗
FerryAI
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process…
★ 41
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.