tonyd2wild/GLM-5.2-NVFP4-KV-4x-DGX-Spark-300kctx-42tok-s
GLM-5.2 744B with a true 4-bit NVFP4 KV cache on 4x DGX Spark: 42 tok/s peak, 317K-token KV pool (+58.6% vs fp8), needle-verified at 250K depth, serving 316K context
⭐ 14
⑂ 0
Python
Apache-2.0
· 2026-07-21推送
14
Watchers
0
贡献者
0
Commits
0
Releases
2
Open Issues
2026-07-21
最近推送
原文
中文
暂无 README