evaluation
101 个项目 · ⭐ 261.2kOpen-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Supercharge Your LLM Application Evaluations 🚀
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
The platform for LLM evaluations and AI agent testing
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
SuperCLUE: 中文通用大模型综合性基准 | A Benchmark for Foundation Models in Chinese
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
An open-source visual programming environment for battle-testing prompts to LLMs.
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
共 101 条 · 第 1 / 6 页