← 返回专题广场
fp8
6 个项目 · ⭐ 3.6k1
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Python
⭐ 3.5k
⑂ 805
Apache-2.0
· 7 小时前推送
7 小时前
最近推送
2
29 天前
最近推送
3
2026-05-16
最近推送
4
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
Python
⭐ 14
⑂ 1
· 2026-07-10推送
2026-07-10
最近推送
5
2026-06-07
最近推送