jiahongsigma/Efficient-LLM-Inference-Serving-Systems
Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
⭐ 20
⑂ 2
Python
MIT
· 11 天前推送
20
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
11 天前
最近推送
原文
中文
暂无 README