TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtim
TensorRT-LLM compiles models into highly optimized NVIDIA kernels with in-flight batching, quantization and multi-GPU tensor parallelism — the reference for squeezing maximum tokens per second from NVIDIA hardware.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtim
TensorRT-LLM has 14.5k stars on GitHub. It has been forked 2.7k times. TensorRT-LLM is written mainly in C++/Python. It has been in active development since 2023. Its main topics are blackwell, cuda, llm-serving, moe.
Read the full guideTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtim
TensorRT-LLM is an open-source project.
Yes. TensorRT-LLM is free and open source — you can use, modify and self-host it.
TensorRT-LLM is written mainly in C++/Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/nvidia-tensorrt-llm.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.