massif-01/vllm_benchmark_block_fp8
Automated Triton w8a8 block FP8 kernel tuning tool for vLLM. Auto-detects model architecture, supports Qwen3-Coder-30B-A3B-Instruct-FP8/DeepSeek-V3/custom models, multi-GPU parallel tuning, and generates optimized kernel configs for quantization.
⭐ 15
⑂ 2
Python
NOASSERTION
· 2026-05-16推送
15
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
2026-05-16
最近推送
原文
中文
暂无 README