evaluation
101 个项目 · ⭐ 261.2kGenerate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
Framework for enhancing LLMs for RAG tasks using fine-tuning.
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
🦞 Official plugin for OpenClaw that exports agent traces to Opik. See and monitor agent behaviour, cost, tokens, errors and more.
Open-source benchmark for browser AI agents on daily tasks.
Dataset and benchmark for RAG on company internal documents.
Resource, Evaluation and Detection Papers for ChatGPT
Harness-oriented agent system framework for production-grade LLM agent applications
A Comprehensive Framework for Building End-to-End Recommendation Systems with State-of-the-Art Models
[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
XRAG: eXamining the Core - Benchmarking Foundational Component Modules in Advanced Retrieval-Augmented Generation
Fiddler Auditor is a tool to evaluate language models.
共 101 条 · 第 3 / 6 页