NVIDIA

NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

⭐ 14.5k ⑂ 2.7k Python NOASSERTION · 5 小时前推送
14.5k
Watchers
0
贡献者
0
Commits
0
Releases
1.5k
Open Issues
5 小时前
最近推送

📦 版本动态

v1.3.0rc24 预发布 10 天前
v1.3.0rc23 预发布 22 天前
v1.3.0rc22 预发布 2026-07-22
v1.3.0rc21 预发布 2026-07-15
v1.3.0rc20 预发布 2026-06-30
原文 中文
暂无 README