tonyd2wild

tonyd2wild/GLM-5.2-NVFP4-KV-4x-DGX-Spark-300kctx-42tok-s

GLM-5.2 744B with a true 4-bit NVFP4 KV cache on 4x DGX Spark: 42 tok/s peak, 317K-token KV pool (+58.6% vs fp8), needle-verified at 250K depth, serving 316K context

⭐ 14 ⑂ 0 Python Apache-2.0 · 2026-07-21推送
14
Watchers
0
贡献者
0
Commits
0
Releases
2
Open Issues
2026-07-21
最近推送
原文 中文
暂无 README