← 返回专题广场
lm-eval
1 个项目 · ⭐ 1021
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
Python
⭐ 102
⑂ 11
NOASSERTION
· 18 小时前推送
18 小时前
最近推送