evaluation
101 个项目 · ⭐ 261.3kA unified evaluation framework for large language models
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
🤗 Evaluate: A library for easily evaluating machine learning models and datasets.
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
Avalanche: an End-to-End Library for Continual Learning based on PyTorch.
:cloud: :rocket: :bar_chart: :chart_with_upwards_trend: Evaluating state of the art in AI
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
(IROS 2020, ECCVW 2020) Official Python Implementation for "3D Multi-Object Tracking: A Baseline and New Evaluation Metrics"
Building blocks for rapid development of GenAI applications
Evaluate your LLM's response with Prometheus and GPT4 💯
共 101 条 · 第 2 / 6 页