jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
⭐ 25
⑂ 1
Python
NOASSERTION
· 18 天前推送
25
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
18 天前
最近推送
原文
中文
暂无 README