llm-inference

76 progetti condividono questo topic GitHub

llm-inference — gpt4all ★77.4kllm-inferenceray — ★43.4kgitleaks — ★28.4kllm-action — ★24.8klitgpt — ★13.6kOpenLLM — ★12.4kopenvino — ★10.6kPowerInfer — ★9.7kBentoML — ★8.7klmdeploy — ★8kplano — ★6.9kflashinfer — ★6.1kkserve — ★5.7kshimmy — ★5.7kcactus — ★5.5kAwesome-LLM-Inference — ★5.4kgpustack — ★5.4ksuperduper — ★5.3klemonade — ★5.2keko — ★4.9kRuVector — ★4.4koptillm — ★4.2kGenerativeAIExamples — ★4.1klorax — ★3.8kspiceai — ★3.1kdistributed-llama — ★3kMedusa — ★2.8kEAGLE — ★2.5knanocoder — ★2.3kaici — ★2.1kneuron-ai — ★2kllama2-webui — ★1.9kbeta9 — ★1.7kreact-native-executorch — ★1.7kxllm — ★1.5kBrowserAI — ★1.4kawesome-ai-web-search — ★1.4kawesome-hacking-lists — ★1.4kLeanCopilot — ★1.3kAtomic-Chat — ★1.2ktiny-vllm — ★963ray★ 43.4kgitleaks★ 28.4kllm-action★ 24.8klitgpt★ 13.6kOpenLLM★ 12.4kopenvino★ 10.6kPowerInfer★ 9.7kBentoML★ 8.7klmdeploy★ 8kplano★ 6.9kflashinfer★ 6.1kkserve★ 5.7kshimmy★ 5.7kcactus★ 5.5kAwesome-LLM-Inference★ 5.4kgpustack★ 5.4ksuperduper★ 5.3klemonade★ 5.2keko★ 4.9kRuVector★ 4.4koptillm★ 4.2kGenerativeAIExamples★ 4.1klorax★ 3.8kspiceai★ 3.1kdistributed-llama★ 3kMedusa★ 2.8kEAGLE★ 2.5knanocoder★ 2.3kaici★ 2.1kneuron-ai★ 2kllama2-webui★ 1.9kbeta9★ 1.7kreact-native-executorch★ 1.7kxllm★ 1.5kBrowserAI★ 1.4kawesome-ai-web-search★ 1.4kawesome-hacking-lists★ 1.4kLeanCopilot★ 1.3kAtomic-Chat★ 1.2ktiny-vllm★ 963

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
gpt4all
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77.4k
ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for…
★ 43.4k
gitleaks
Find secrets with Gitleaks 🔑
★ 28.4k
llm-action
★ 24.8k
litgpt
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
★ 13.6k
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12.4k
openvino
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
★ 10.6k
PowerInfer
High-speed Large Language Model Serving for Local Deployment
★ 9.7k
BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model…
★ 8.7k
lmdeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
★ 8k
plano
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent…
★ 6.9k
flashinfer
FlashInfer: Kernel Library for LLM Serving
★ 6.1k
kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework…
★ 5.7k
shimmy
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No…
★ 5.7k
cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
★ 5.5k
Awesome-LLM-Inference
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4,…
★ 5.4k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.4k
superduper
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5.3k
lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and…
★ 5.2k
eko
Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai
★ 4.9k
RuVector
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
★ 4.4k
optillm
Optimizing inference proxy for LLMs
★ 4.2k
GenerativeAIExamples
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
★ 4.1k
lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
★ 3.8k
spiceai
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query,…
★ 3.1k
distributed-llama
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More…
★ 3k
Medusa
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8k
EAGLE
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2.5k
nanocoder
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own…
★ 2.3k
aici
AICI: Prompts as (Wasm) Programs
★ 2.1k
neuron-ai
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect…
★ 2k
llama2-webui
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper`…
★ 1.9k
beta9
Ultrafast serverless GPU inference, sandboxes, and background jobs
★ 1.7k
react-native-executorch
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1.7k
xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators.…
★ 1.5k
BrowserAI
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser
★ 1.4k
awesome-ai-web-search
List of software that allows searching the web with the assistance of AI:…
★ 1.4k
awesome-hacking-lists
A curated collection of top-tier penetration testing tools and productivity utilities across multiple…
★ 1.4k
LeanCopilot
LLMs as Copilots for Theorem Proving in Lean
★ 1.3k
Atomic-Chat
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your…
★ 1.2k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 963
ZhiLight
A highly optimized LLM inference acceleration engine for Llama and its variants.
★ 907
llama3.java
Llama 3+ inference in pure Java
★ 816
blast
Open-source VMs-as-a-service
★ 778
LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing
LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for…
★ 731
llmflows
LLMFlows - Simple, Explicit and Transparent LLM Apps
★ 707
atlas
Pure Rust Inference Engine
★ 618
MiniSearch
Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo:…
★ 579
LLM-FineTuning-Large-Language-Models
LLM (Large Language Model) FineTuning
★ 576
LLM-Hub
Local AI Assistant on your phone
★ 518
krasis
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM…
★ 494
edsl
Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market…
★ 484
chipper
✨ AI interface for tinkerers (Ollama, Haystack RAG, Python)
★ 483
SwiftInfer
Efficient AI Inference & Serving
★ 478
taOS
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default,…
★ 464
incognide
Explore the unknown, build the future, own your data.
★ 459
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 451
sagify
LLMs and Machine Learning done easily
★ 442
Context-Engine
Context-Engine MCP - Agentic Context Compression Suite
★ 402
Star-Attention
Efficient LLM Inference over Long Sequences
★ 391
embedding_studio
Embedding Studio is a framework which allows you transform your Vector Database into a feature-rich Search…
★ 382
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models,…
★ 381
HackBot
AI-powered cybersecurity chatbot designed to provide helpful and accurate answers to your…
★ 357
syncode
Efficient and general syntactical decoding for Large Language Models
★ 338
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 330
MoE-Infinity
PyTorch library for cost-effective, fast and easy serving of MoE models.
★ 328
monocle
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI…
★ 318
leanctx
Drop-in prompt compression for production LLM apps. Cut your token bill 40-60% without changing your code.…
★ 316
deepseek.cpp
CPU inference for the DeepSeek family of large language models in C++
★ 316
picollm
On-device LLM Inference Powered by X-Bit Quantization
★ 315
ecologits
🌱 EcoLogits tracks the energy consumption and environmental footprint of using generative AI models…
★ 307
pmetal
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving,…
★ 306
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 128
FerryAI
Native AI inference for PHP 8.3+ - run ONNX, GGUF (llama.cpp) and RubixML models directly in your PHP process…
★ 56 · GitHub ↗
llmtrace
Zero-code LLM security & observability proxy. Real-time prompt injection detection, PII scanning, and cost…
★ 52 · GitHub ↗
llmff
Bounded, inspectable LLM inference pipelines from declared YAML — runs offline against Ollama or any…
★ 45 · GitHub ↗
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.