quantization

51 projetos partilham este topic do GitHub

quantization — LlamaFactory ★73.6kquantizationfaster-whisper — ★24.6kChinese-LLaMA-Alpaca — ★18.9kQbot — ★18.2kturbovec — ★14.5kbitsandbytes — ★8.4kCTranslate2 — ★4.6knunchaku — ★3.9kSageAttention — ★3.5koptimum — ★3.5kPretrained-Language-Model — ★3.2kneural-compressor — ★2.7kaimet — ★2.7kxTuring — ★2.7kAwesome-Model-Quantization — ★2.4kAI-Engineering.academy — ★2.4kmixtral-offloading — ★2.3kai-engineering-interview-questions — ★2.3kppq — ★1.8kpicolm — ★1.7krwkv.cpp — ★1.6kmodel-optimization — ★1.6kbrevitas — ★1.6kauto-round — ★1.5kAngelSlim — ★1.5kgpu_poor — ★1.4kgeti — ★1.3kGPTQModel — ★1.2knncf — ★1.2kz80ai — ★1.1kMindPipe — ★1kTinyChatEngine — ★959OmniQuant — ★905Awesome-Quantization-Papers — ★838LightCompress — ★736SqueezeLLM — ★722hailo_model_zoo — ★684SINQ — ★627tidy — ★586minigpt4.cpp — ★573QeRL — ★511faster-whisper★ 24.6kChinese-LLaMA-Alpaca★ 18.9kQbot★ 18.2kturbovec★ 14.5kbitsandbytes★ 8.4kCTranslate2★ 4.6knunchaku★ 3.9kSageAttention★ 3.5koptimum★ 3.5kPretrained-Language-Mode…★ 3.2kneural-compressor★ 2.7kaimet★ 2.7kxTuring★ 2.7kAwesome-Model-Quantizati…★ 2.4kAI-Engineering.academy★ 2.4kmixtral-offloading★ 2.3kai-engineering-interview…★ 2.3kppq★ 1.8kpicolm★ 1.7krwkv.cpp★ 1.6kmodel-optimization★ 1.6kbrevitas★ 1.6kauto-round★ 1.5kAngelSlim★ 1.5kgpu_poor★ 1.4kgeti★ 1.3kGPTQModel★ 1.2knncf★ 1.2kz80ai★ 1.1kMindPipe★ 1kTinyChatEngine★ 959OmniQuant★ 905Awesome-Quantization-Pap…★ 838LightCompress★ 736SqueezeLLM★ 722hailo_model_zoo★ 684SINQ★ 627tidy★ 586minigpt4.cpp★ 573QeRL★ 511

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 73.6k
faster-whisper
Faster Whisper transcription with CTranslate2
★ 24.6k
Chinese-LLaMA-Alpaca
中文LLaMA&Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA & Alpaca LLMs)
★ 18.9k
Qbot
[🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment…
★ 18.2k
turbovec
A vector index built on TurboQuant, written in Rust with Python bindings
★ 14.5k
bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
★ 8.4k
CTranslate2
Fast inference engine for Transformer models
★ 4.6k
nunchaku
[ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
★ 3.9k
SageAttention
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to…
★ 3.5k
optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with…
★ 3.5k
Pretrained-Language-Model
Pretrained language model and its related optimization techniques developed by Huawei Noah's Ark Lab.
★ 3.2k
neural-compressor
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression…
★ 2.7k
aimet
AIMET is a library that provides advanced quantization and compression techniques for trained neural network…
★ 2.7k
xTuring
Build, personalize and control your own LLMs. From data pre-processing to fine-tuning, xTuring provides an…
★ 2.7k
Awesome-Model-Quantization
A list of papers, docs, codes about model quantization. This repo is aimed to provide the info for model…
★ 2.4k
AI-Engineering.academy
Mastering Applied AI, One Concept at a Time
★ 2.4k
mixtral-offloading
Run Mixtral-8x7B models in Colab or consumer desktops
★ 2.3k
ai-engineering-interview-questions
Your Cheat Sheet for AI Engineering Interview – Questions and Answers.
★ 2.3k
ppq
PPL Quantization Tool (PPQ) is a powerful offline neural network quantization tool.
★ 1.8k
picolm
Run a 1-billion parameter LLM on a $10 board with 256MB RAM
★ 1.7k
rwkv.cpp
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
★ 1.6k
model-optimization
A toolkit to optimize ML models for deployment for Keras and TensorFlow, including quantization and pruning.
★ 1.6k
brevitas
Brevitas: neural network quantization in PyTorch
★ 1.6k
auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA,…
★ 1.5k
AngelSlim
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1.5k
gpu_poor
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
★ 1.4k
geti
Build computer vision models in a fraction of the time and with less data.
★ 1.3k
GPTQModel
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and…
★ 1.2k
nncf
Neural Network Compression Framework for enhanced OpenVINO™ inference
★ 1.2k
z80ai
Z80-μLM is a 2-bit quantized language model small enough to run on an 8-bit Z80 processor. Train…
★ 1.1k
MindPipe
A powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
★ 1k
TinyChatEngine
TinyChatEngine: On-Device LLM Inference Library
★ 959
OmniQuant
[ICLR2024 spotlight] OmniQuant is a simple and powerful quantization technique for LLMs.
★ 905
Awesome-Quantization-Papers
List of papers related to neural network quantization in recent AI conferences and journals.
★ 838
LightCompress
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video…
★ 736
SqueezeLLM
[ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization
★ 722
hailo_model_zoo
The Hailo Model Zoo includes pre-trained models and a full building and evaluation environment
★ 684
SINQ
Welcome to the official repository of SINQ! A novel, fast and high-quality quantization method designed to…
★ 627
tidy
Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art…
★ 586
minigpt4.cpp
Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)
★ 573
QeRL
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
★ 511
KVQuant
[NeurIPS 2024] KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
★ 430
PiSSA
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models(NeurIPS 2024…
★ 429
KIVI
[ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
★ 423
quick-start-guide-to-llms
The Official Repo for "Quick Start Guide to Large Language Models"
★ 393
q-diffusion
[ICCV 2023] Q-Diffusion: Quantizing Diffusion Models.
★ 378
KVSplit
Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache…
★ 361
BitNet-Transformers
0️⃣1️⃣🤗 BitNet-Transformers: Huggingface Transformers Implementation of "BitNet: Scaling 1-bit…
★ 316
picollm
On-device LLM Inference Powered by X-Bit Quantization
★ 315
pmetal
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving,…
★ 306
sndr_core_engine
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere…
★ 128
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.