evaluation-metrics
17 个项目 · ⭐ 38.5kLighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
Evaluate your speech-to-text system with similarity measures such as word error rate (WER)
Data-Driven Evaluation for LLM-Powered Applications
Benchmark diffusion models faster. Automate evals, seeds, and metrics for reproducible results.
RAG evaluation without the need for "golden answers"
[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Language AI Engineering Lab, a place where you can deeply understand and build modern Language AI systems, from fundamentals to production.
한국어 STT 출력의 CER, WER, CRR, 키워드·개체명·코퍼스 평가를 제공하는 Python 패키지
Evals framework for Information Retrieval Systems