jiahongsigma

jiahongsigma/Efficient-LLM-Inference-Serving-Systems

Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.

⭐ 20 ⑂ 2 Python MIT · 11 天前推送
20
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
11 天前
最近推送
原文 中文
暂无 README