llm-evaluation
52 个项目 · ⭐ 205.5kA test runner for agentskills.io-style AI agent skills
Open-source benchmark for browser AI agents on daily tasks.
Dataset and benchmark for RAG on company internal documents.
Data-Driven Evaluation for LLM-Powered Applications
Build, Improve Performance, and Productionize your AI Application
Engineering deterministic, production-grade systems around non-deterministic LLMs — FSM, durable execution, retries, DAGs, agent runtimes, model routing, edge inference, RAG, memory, multi-agent orchestration, security, and observability. 14 runnable proof-of-concept phases.
Rank LLMs, RAG systems, and prompts using automated head-to-head evaluation
A desktop application for comparing outputs from different Large Language Models (LLMs).
共 52 条 · 第 2 / 3 页