evaluation-framework
19 个项目 · ⭐ 69.9kA framework for few-shot evaluation of language models.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI
Data-Driven Evaluation for LLM-Powered Applications
MedEvalKit: A Unified Medical Evaluation Framework
WritingBench: A Comprehensive Benchmark for Generative Writing
The official implementation of ECCV'24 paper "To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now". This work introduces one fast and effective attack method to evaluate the harmful-content generation ability of safety-driven unlearned diffusion models.
Framework to evaluate Trajectory Classification Algorithms
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)