OnlyTerp/turboquant
First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
⭐ 78
⑂ 11
Python
MIT
· 2026-05-25推送
78
Watchers
0
贡献者
0
Commits
0
Releases
1
Open Issues
2026-05-25
最近推送
原文
中文
暂无 README