ianarawjo

ianarawjo/evalstats

Statistical analysis for LLM evaluations, from model and prompt comparisons to inference resilient to LLM judge bias, including at small sample sizes. All defaults battle-tested in Monte Carlo simulations.

⭐ 115 ⑂ 2 Python NOASSERTION · 4 天前推送
115
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
4 天前
最近推送
原文 中文
暂无 README