cuda

90 progetti condividono questo topic GitHub

cuda — vllm ★87.6kcudavoicebox — ★47.4ksglang — ★30.9kinstant-ngp — ★17.5kburn — ★15.7kkaldi — ★15.4kTensorRT-LLM — ★14.2kOpen3D — ★13.8kGPU-Puzzles — ★12.3kLMCache — ★10.9kcutlass — ★10.2kcog — ★9.4koneflow — ★9.4kcatboost — ★9kgocv — ★7.5kflashinfer — ★6.1kchainer — ★5.9kgpustack — ★5.4kcuml — ★5.2knccl — ★4.9kCTranslate2 — ★4.6kTengine — ★4.5ktiny-cuda-nn — ★4.5kiree — ★3.9kSageAttention — ★3.5kLichtFeld-Studio — ★3.5kTransformerEngine — ★3.5kjittor — ★3.2khow-to-optim-algorithm-in-cuda — ★3.2klc0 — ★3.2kheavydb — ★3.1kTensorRT — ★3kramalama — ★3kMinkowskiEngine — ★2.9kskills — ★2.7kCV-CUDA — ★2.7kCVprojects — ★2.6ktorchrec — ★2.6knode-llama-cpp — ★2.1kpykeen — ★2kdeepmd-kit — ★2kvoicebox★ 47.4ksglang★ 30.9kinstant-ngp★ 17.5kburn★ 15.7kkaldi★ 15.4kTensorRT-LLM★ 14.2kOpen3D★ 13.8kGPU-Puzzles★ 12.3kLMCache★ 10.9kcutlass★ 10.2kcog★ 9.4koneflow★ 9.4kcatboost★ 9kgocv★ 7.5kflashinfer★ 6.1kchainer★ 5.9kgpustack★ 5.4kcuml★ 5.2knccl★ 4.9kCTranslate2★ 4.6kTengine★ 4.5ktiny-cuda-nn★ 4.5kiree★ 3.9kSageAttention★ 3.5kLichtFeld-Studio★ 3.5kTransformerEngine★ 3.5kjittor★ 3.2khow-to-optim-algorithm-i…★ 3.2klc0★ 3.2kheavydb★ 3.1kTensorRT★ 3kramalama★ 3kMinkowskiEngine★ 2.9kskills★ 2.7kCV-CUDA★ 2.7kCVprojects★ 2.6ktorchrec★ 2.6knode-llama-cpp★ 2.1kpykeen★ 2kdeepmd-kit★ 2k

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 87.6k
voicebox
The open-source AI voice studio. Clone, dictate, create.
★ 47.4k
sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 30.9k
instant-ngp
Instant neural graphics primitives: lightning fast NeRF and more
★ 17.5k
burn
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility,…
★ 15.7k
kaldi
kaldi-asr/kaldi is the official location of the Kaldi project.
★ 15.4k
TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and…
★ 14.2k
Open3D
Open3D: A Modern Library for 3D Data Processing
★ 13.8k
GPU-Puzzles
Solve puzzles. Learn CUDA.
★ 12.3k
LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 10.9k
cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10.2k
cog
Containers for machine learning
★ 9.4k
oneflow
OneFlow is a deep learning framework designed to be user-friendly, scalable and efficient.
★ 9.4k
catboost
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking,…
★ 9k
gocv
Go package for computer vision using OpenCV 4 and beyond. Includes support for DNN, CUDA, OpenCV Contrib, and…
★ 7.5k
flashinfer
FlashInfer: Kernel Library for LLM Serving
★ 6.1k
chainer
A flexible framework of neural networks for deep learning
★ 5.9k
gpustack
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU…
★ 5.4k
cuml
cuML - RAPIDS Machine Learning Library
★ 5.2k
nccl
Optimized primitives for collective multi-GPU communication
★ 4.9k
CTranslate2
Fast inference engine for Transformer models
★ 4.6k
Tengine
Tengine is a lite, high performance, modular inference engine for embedded device
★ 4.5k
tiny-cuda-nn
Lightning fast C++/CUDA neural network framework
★ 4.5k
iree
A retargetable MLIR-based machine learning compiler and runtime toolkit.
★ 3.9k
SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to…
★ 3.5k
LichtFeld-Studio
Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
★ 3.5k
TransformerEngine
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point…
★ 3.5k
jittor
Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.
★ 3.2k
how-to-optim-algorithm-in-cuda
how to optimize some algorithm in cuda.
★ 3.2k
lc0
Open source neural network chess engine with GPU acceleration and broad hardware support.
★ 3.2k
heavydb
HeavyDB (formerly MapD/OmniSciDB)
★ 3.1k
TensorRT
PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT
★ 3k
ramalama
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and…
★ 3k
MinkowskiEngine
Minkowski Engine is an auto-diff neural network library for high-dimensional sparse tensors
★ 2.9k
skills
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical…
★ 2.7k
CV-CUDA
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
★ 2.7k
CVprojects
computer vision projects | 计算机视觉相关好玩的AI项目(Python、C++、embedded system)
★ 2.6k
torchrec
Pytorch domain library for recommendation systems
★ 2.6k
node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model…
★ 2.1k
pykeen
🤖 A Python library for learning and evaluating knowledge graph embeddings
★ 2k
deepmd-kit
A deep learning package for many-body potential energy representation and molecular dynamics
★ 2k
onediff
OneDiff: An out-of-the-box acceleration library for diffusion models.
★ 2k
dfdx
Deep learning in Rust, with shape checked tensors and neural networks
★ 1.9k
sonar
Large-scale LLM inference engine
★ 1.8k
ppq
PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool.
★ 1.8k
awesome-yolo-object-detection
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object…
★ 1.8k
curobo
CUDA Accelerated Robot Library
★ 1.7k
beta9
Ultrafast serverless GPU inference, sandboxes, and background jobs
★ 1.7k
tt-metal
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
★ 1.6k
gpu-hot
🔥 Real-time NVIDIA GPU dashboard
★ 1.6k
3d-ken-burns
an implementation of 3D Ken Burns Effect from a Single Image using PyTorch
★ 1.6k
uccl
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL…
★ 1.5k
Chatterbox-TTS-Server
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API…
★ 1.4k
stable-fast
https://wavespeed.ai/ Best inference performance optimization framework for HuggingFace Diffusers on NVIDIA…
★ 1.3k
InferenceX
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 -…
★ 1.3k
cupoch
Robotics with GPU computing
★ 1.1k
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 963
ZhiLight
A highly optimized LLM inference acceleration engine for Llama and its variants.
★ 907
gprMax
gprMax is open source software that simulates electromagnetic wave propagation using the Finite-Difference…
★ 864
UniLab
UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
★ 850
Savant
Python Computer Vision & Video Analytics Framework With Batteries Included
★ 839
GPUMD
Graphics Processing Units Molecular Dynamics
★ 818
caer
High-performance Vision library in Python. Scale your research, not boilerplate.
★ 812
awesome-llm-and-aigc
🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language…
★ 811
surogate
Training/Fine-tuning at the speed of light
★ 806
ServerlessLLM
Serverless LLM Serving for Everyone.
★ 695
chatterbox-tts-api
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned…
★ 630
atlas
Pure Rust Inference Engine
★ 618
vins-application
VINS-Fusion, VINS-Fisheye, OpenVINS, EnVIO, ROVIO, S-MSCKF, ORB-SLAM2, NVIDIA Elbrus application of different…
★ 610
attorch
A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.
★ 604
SAM3DBody-cpp
Real-time 3D full-body reconstruction from a single camera, Multiperson BVH output, Pure C++ runtime, ONNX +…
★ 604
llm_training_handbook
An open collection of methodologies to help with successful training of large language models.
★ 566
radarsimpy
Radar Simulator built with Python and C++
★ 561
willow-inference-server
Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and…
★ 510
large_language_model_training_playbook
An open collection of implementation tips, tricks and resources for training large language models
★ 503
popsift
PopSift is an implementation of the SIFT algorithm in CUDA.
★ 499
cucim
cuCIM - RAPIDS GPU-accelerated image processing library
★ 465
hoomd-blue
Molecular dynamics and Monte Carlo soft matter simulation on GPUs.
★ 444
dynamicfusion
Implementation of Newcombe et al. CVPR 2015 DynamicFusion paper
★ 413
splatad
SplatAD: Real-Time Lidar and Camera Rendering with 3D Gaussian Splatting for Autonomous Driving
★ 407
MFC
Exascale multiphase flow solver — 2025 Gordon Bell Prize Finalist | 200T grid points on 43K+ GPUs
★ 388
Dia-TTS-Server
Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints…
★ 352
swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with…
★ 330
dynamic-occupancy-grid-map
Implementation of "A Random Finite Set Approach for Dynamic Occupancy Grid Maps with Real-Time Application"
★ 310
MPPI-Generic
Templated C++/CUDA implementation of Model Predictive Path Integral Control (MPPI)
★ 301
Inline-Studio
AI filmmaking on a node canvas. Build your whole visual pipeline from moodboard to final cut. Generate on…
★ 180
self-hosted-ai-stack
Deploy a complete self-hosted AI stack with Docker Compose: Ollama, LiteLLM, AnythingLLM, Whisper,…
★ 129
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 128
dgx-spark-inference-stack
Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your…
★ 50 · GitHub ↗
trellis2.c
Generate textured, segmented and rigged GLB assets entirely on your own GPU.
★ 48 · GitHub ↗
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.