dipampaul17

dipampaul17/KVSplit

Run larger LLMs with longer contexts on Apple Silicon by using differentiated precision for KV cache quantization. KVSplit enables 8-bit keys & 4-bit values, reducing memory by 59% with <1% quality loss. Includes benchmarking, visualization, and one-command setup. Optimized for M1/M2/M3 Macs with Metal support.

⭐ 362 ⑂ 13 Python NOASSERTION · 2025-05-21推送
362
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
2025-05-21
最近推送
原文 中文
暂无 README