← 返回专题广场
llm-evaluation
52 个项目 · ⭐ 205.5k1
4 小时前
最近推送
2
16 小时前
最近推送
3
5 小时前
最近推送
4
1 天前
最近推送
6
1 天前
最近推送
8
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
TypeScript
⭐ 6.1k
⑂ 658
Apache-2.0
· 6 天前推送
6 天前
最近推送
9
🐢 Open-Source Evaluation & Testing library for LLM Agents
Python
⭐ 5.8k
⑂ 522
Apache-2.0
· 2 天前推送
2 天前
最近推送
10
5 小时前
最近推送
11
1 天前
最近推送
12
2026-04-22
最近推送
13
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
TypeScript
⭐ 5.1k
⑂ 432
NOASSERTION
· 7 小时前推送
7 小时前
最近推送
14
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Python
⭐ 4.4k
⑂ 645
NOASSERTION
· 12 小时前推送
12 小时前
最近推送
15
Evaluation and Tracking for LLM Experiments and AI Agents
Python
⭐ 3.5k
⑂ 329
MIT
· 1 天前推送
1 天前
最近推送
16
17 天前
最近推送
17
6 小时前
最近推送
18
A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers identify and mitigate potential jailbreaks in their LLM APIs.
Jupyter Notebook
⭐ 1.6k
⑂ 215
Apache-2.0
· 2026-02-07推送
2026-02-07
最近推送
19
16 小时前
最近推送
20
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Python
⭐ 1.1k
⑂ 96
Apache-2.0
· 5 天前推送
5 天前
最近推送
共 52 条 · 第 1 / 3 页