← 返回专题广场
llm-as-a-judge
4 个项目 · ⭐ 1.9k1
Evaluate your LLM's response with Prometheus and GPT4 💯
Python
⭐ 1.1k
⑂ 68
Apache-2.0
· 2025-04-25推送
2025-04-25
最近推送
2
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
Python
⭐ 741
⑂ 75
Apache-2.0
· 15 小时前推送
15 小时前
最近推送
3
Evals framework for Information Retrieval Systems
Python
⭐ 18
⑂ 3
MIT
· 2026-07-19推送
2026-07-19
最近推送
4
⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.
Python
⭐ 15
⑂ 2
Apache-2.0
· 2026-05-29推送
2026-05-29
最近推送