gpu

93 progetti condividono questo topic GitHub

gpu — pytorch ★102kgpuDeepSpeed — ★42.8kfastai — ★28.1kWebGL-Fluid-Simulation — ★16.5kOpen3D — ★13.8ktvm — ★13.6kserver — ★10.9kskypilot — ★10.4kcutlass — ★10.2kcatboost — ★9kImageAI — ★8.9kh2o-3 — ★7.5kflashinfer — ★6.1kwhichllm — ★6kchainer — ★5.9kDALI — ★5.7kcuml — ★5.2kFluidX3D — ★5.2klemonade — ★5.2kkoharu — ★5kpytorch-forecasting — ★5knccl — ★4.9kexecutorch — ★4.8kMegEngine — ★4.8ktiny-cuda-nn — ★4.5kdeepflow — ★4.2kllm-d — ★3.9ktvm-cn — ★3.9kml-workspace — ★3.5kTransformerEngine — ★3.5kjittor — ★3.2klc0 — ★3.2kchitu — ★3.1kheavydb — ★3.1kleptonai — ★2.8kdifftaichi — ★2.7kCV-CUDA — ★2.7ktorchrec — ★2.6kdeepdetect — ★2.6khyperlearn — ★2.5kgpupixel — ★2.3kDeepSpeed★ 42.8kfastai★ 28.1kWebGL-Fluid-Simulation★ 16.5kOpen3D★ 13.8ktvm★ 13.6kserver★ 10.9kskypilot★ 10.4kcutlass★ 10.2kcatboost★ 9kImageAI★ 8.9kh2o-3★ 7.5kflashinfer★ 6.1kwhichllm★ 6kchainer★ 5.9kDALI★ 5.7kcuml★ 5.2kFluidX3D★ 5.2klemonade★ 5.2kkoharu★ 5kpytorch-forecasting★ 5knccl★ 4.9kexecutorch★ 4.8kMegEngine★ 4.8ktiny-cuda-nn★ 4.5kdeepflow★ 4.2kllm-d★ 3.9ktvm-cn★ 3.9kml-workspace★ 3.5kTransformerEngine★ 3.5kjittor★ 3.2klc0★ 3.2kchitu★ 3.1kheavydb★ 3.1kleptonai★ 2.8kdifftaichi★ 2.7kCV-CUDA★ 2.7ktorchrec★ 2.6kdeepdetect★ 2.6khyperlearn★ 2.5kgpupixel★ 2.3k

Le linee collegano membri che sono misurabilmente correlati tra loro. La dimensione dei punti riflette le stelle.

🧬 Membri
pytorch
Tensors and Dynamic neural networks in Python with strong GPU acceleration
★ 102k
DeepSpeed
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy,…
★ 42.8k
fastai
The fastai deep learning library
★ 28.1k
WebGL-Fluid-Simulation
Play with fluids in your browser (works even on mobile)
★ 16.5k
Open3D
Open3D: A Modern Library for 3D Data Processing
★ 13.8k
tvm
Open Machine Learning Compiler Framework
★ 13.6k
server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
★ 10.9k
skypilot
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer,…
★ 10.4k
cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
★ 10.2k
catboost
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking,…
★ 9k
ImageAI
A python library built to empower developers to build applications and systems with self-contained Computer…
★ 8.9k
h2o-3
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient…
★ 7.5k
flashinfer
FlashInfer: Kernel Library for LLM Serving
★ 6.1k
whichllm
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware…
★ 6k
chainer
A flexible framework of neural networks for deep learning
★ 5.9k
DALI
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data…
★ 5.7k
cuml
cuML - RAPIDS Machine Learning Library
★ 5.2k
FluidX3D
The fastest and most memory efficient lattice Boltzmann CFD software, running on all GPUs and CPUs via…
★ 5.2k
lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and…
★ 5.2k
koharu
ML-powered manga translator, written in Rust.
★ 5k
pytorch-forecasting
Time series forecasting with PyTorch
★ 5k
nccl
Optimized primitives for collective multi-GPU communication
★ 4.9k
executorch
On-device AI across mobile, embedded and edge for PyTorch
★ 4.8k
MegEngine
MegEngine 是一个快速、可拓展、易于使用且支持自动求导的深度学习框架
★ 4.8k
tiny-cuda-nn
Lightning fast C++/CUDA neural network framework
★ 4.5k
deepflow
eBPF Observability - Distributed Tracing and Profiling
★ 4.2k
llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
★ 3.9k
tvm-cn
TVM Documentation in Chinese Simplified / TVM 中文文档
★ 3.9k
ml-workspace
🛠 All-in-one web-based IDE specialized for machine learning and data science.
★ 3.5k
TransformerEngine
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point…
★ 3.5k
jittor
Jittor is a high-performance deep learning framework based on JIT compiling and meta-operators.
★ 3.2k
lc0
Open source neural network chess engine with GPU acceleration and broad hardware support.
★ 3.2k
chitu
High-performance inference framework for large language models, focusing on efficiency, flexibility, and…
★ 3.1k
heavydb
HeavyDB (formerly MapD/OmniSciDB)
★ 3.1k
leptonai
A Pythonic framework to simplify AI service building
★ 2.8k
difftaichi
10 differentiable physical simulators built with Taichi differentiable programming (DiffTaichi, ICLR 2020)
★ 2.7k
CV-CUDA
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
★ 2.7k
torchrec
Pytorch domain library for recommendation systems
★ 2.6k
deepdetect
Deep Learning Server and CLI for Torch and TensorRT
★ 2.6k
hyperlearn
2-2000x faster ML algos, 50% less memory usage, works on all hardware - new and old.
★ 2.5k
gpupixel
Real-time image filter engine based on GPU
★ 2.3k
openlake
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
★ 2.3k
dstack
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and…
★ 2.2k
trainer
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
★ 2.2k
node-llama-cpp
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model…
★ 2.1k
blazingsql
BlazingSQL is a lightweight, GPU accelerated, SQL engine for Python. Built on RAPIDS cuDF.
★ 2k
dfdx
Deep learning in Rust, with shape checked tensors and neural networks
★ 1.9k
beta9
Ultrafast serverless GPU inference, sandboxes, and background jobs
★ 1.7k
glim
GLIM: versatile and extensible point cloud-based 3D localization and mapping framework
★ 1.7k
tt-metal
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
★ 1.6k
gpu-hot
🔥 Real-time NVIDIA GPU dashboard
★ 1.6k
gpu-io
A GPU-accelerated computing library for running physics simulations and other GPGPU computations in a web…
★ 1.5k
uccl
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL…
★ 1.5k
isaac_ros_visual_slam
Visual SLAM/odometry package based on NVIDIA-accelerated cuVSLAM
★ 1.4k
gpu_poor
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
★ 1.4k
modal-examples
Examples of programs built using Modal
★ 1.2k
evotorch
Advanced evolutionary computation library built directly on top of PyTorch, created at NNAISENSE.
★ 1.1k
webgl-wind
Wind power visualization with WebGL particles
★ 1.1k
cupoch
Robotics with GPU computing
★ 1.1k
gdrl
Grokking Deep Reinforcement Learning
★ 1k
femtoGPT
Pure Rust implementation of a minimal Generative Pretrained Transformer
★ 936
gprMax
gprMax is open source software that simulates electromagnetic wave propagation using the Finite-Difference…
★ 864
GPUMD
Graphics Processing Units Molecular Dynamics
★ 818
caer
High-performance Vision library in Python. Scale your research, not boilerplate.
★ 812
can-i-finetune-this
Estimate whether a Hugging Face model fits and fine-tunes on your local GPU.
★ 793
isaac_ros_nvblox
NVIDIA-accelerated 3D scene reconstruction and Nav2 local costmap provider using nvblox
★ 730
horizon
GPU-accelerated terminal board that puts all your sessions on an infinite canvas
★ 673
DFloat11
DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference
★ 652
curl
CURL: Contrastive Unsupervised Representation Learning for Sample-Efficient Reinforcement Learning
★ 605
VerletIntegration
A real-time particle simulation that uses Verlet Integration
★ 582
neurokernel
Neurokernel Project
★ 565
popsift
PopSift is an implementation of the SIFT algorithm in CUDA.
★ 499
zinc
Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon
★ 489
blub
3D fluid simulation experiments in Rust, using WebGPU-rs (WIP)
★ 483
isaac_ros_pose_estimation
Deep learned, NVIDIA-accelerated 3D object pose estimation
★ 480
warpx
WarpX is an advanced Particle-In-Cell code.
★ 473
awsome-distributed-ai
Collection of best practices, reference architectures, model training examples and utilities to train large…
★ 468
cucim
cuCIM - RAPIDS GPU-accelerated image processing library
★ 465
JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs…
★ 451
hoomd-blue
Molecular dynamics and Monte Carlo soft matter simulation on GPUs.
★ 444
Awesome-Distributed-Deep-Learning
A curated list of awesome Distributed Deep Learning resources.
★ 442
rag-chatbot
RAG (Retrieval-augmented generation) ChatBot that provides answers based on contextual information extracted…
★ 435
falldetection_openpifpaf
Fall Detection using OpenPifPaf's Human Pose Estimation model
★ 430
WaterBall
Fluid simulation on a sphere🌏
★ 394
tty7
A terminal workbench in pure Rust: shells, persistent sessions, SSH, coding agents. GPU-rendered on Zed's…
★ 389 · GitHub ↗
MFC
Exascale multiphase flow solver — 2025 Gordon Bell Prize Finalist | 200T grid points on 43K+ GPUs
★ 388
knowhere
Vector search engine inside Milvus, integrating FAISS, HNSW, DiskANN.
★ 372
SPH_Taichi
A high-performance implementation of SPH in Taichi.
★ 320
isaac_ros_common
Common utilities, packages, scripts, and testing infrastructure for Isaac ROS packages.
★ 311
mppi_numba
A GPU implementation of Model Predictive Path Integral (MPPI) control that uses a probabilistic…
★ 310
clearml-agent
ClearML Agent - MLOps/LLMOps made easy. MLOps/LLMOps scheduler & orchestration solution
★ 307
attyx
GPU accelerated terminal for agentic workflows
★ 236
agentfm-core
AgentFM is a peer-to-peer network that turns everyday computers into a decentralized AI supercomputer.…
★ 129
🔗 Famiglie affini

Misurato dai temi di GitHub condivisi da entrambi i progetti, ponderato in base a quanto è raro ciascun tema.