benchmark
142 个项目 · ⭐ 197.5kOpen-source benchmark for browser AI agents on daily tasks.
Dataset and benchmark for RAG on company internal documents.
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
Benchmarking Legal Knowledge of Large Language Models
A universal flight control tuning framework
Gym Electric Motor (GEM): An OpenAI Gym Environment for Electric Motors
LLM Benchmark for Throughput via Ollama (Local LLMs)
Framework for benchmarking vector search engines
Official code repository of < CBGBench: Fill in the Blank of Protein-Molecule Complex Binding Graph >
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
共 142 条 · 第 3 / 8 页