aiha-lab

aiha-lab/tangram

A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM

⭐ 13 ⑂ 1 Python Apache-2.0 · 10 天前推送
13
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
10 天前
最近推送
原文 中文
暂无 README