← 返回专题广场
llm-evaluation
52 个项目 · ⭐ 205.5k41
Python SDK for experimenting, testing, evaluating & monitoring LLM-powered applications - Parea AI (YC S23)
Python
⭐ 82
⑂ 13
Apache-2.0
· 2025-02-13推送
2025-02-13
最近推送
43
2026-06-29
最近推送
44
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
Python
⭐ 44
⑂ 12
Apache-2.0
· 6 小时前推送
6 小时前
最近推送
45
[MM 2025] A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
Python
⭐ 44
⑂ 4
Apache-2.0
· 2026-01-08推送
2026-01-08
最近推送
46
ICLR 2026 | ChemEval: 4-level, 13-dimension, 62-task text/multimodal chemistry benchmark for evaluating LLMs and MLLMs.
Python
⭐ 34
⑂ 1
NOASSERTION
· 2026-06-09推送
2026-06-09
最近推送
47
2026-06-27
最近推送
49
7 小时前
最近推送
50
面向小团队的一站式 NLP 小模型平台:数据集版本管理 · 训练 · 评估 · 在线部署 · Badcase 闭环,内置大模型 Prompt 盲测评测(人工 + AI 双指标)。
Python
⭐ 18
⑂ 2
Apache-2.0
· 2026-06-22推送
2026-06-22
最近推送
51
6 小时前
最近推送
52
Runtime detector for reward hacking and misalignment in LLM agents (89.7% F1 on 5,391 trajectories).
Python
⭐ 12
⑂ 1
Apache-2.0
· 26 天前推送
26 天前
最近推送
共 52 条 · 第 3 / 3 页