benchmark
142 个项目 · ⭐ 197.5kOSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结
Datasets collection and preprocessings framework for NLP extreme multitask learning
WritingBench: A Comprehensive Benchmark for Generative Writing
[ACL 2024] User-friendly evaluation framework: Eval Suite & Benchmarks: UHGEval, HaluEval, HalluQA, etc.
[NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis
[ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
Ollama based Benchmark with detail I/O token per second. Python with Deepseek R1 example.
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Automated system for LLM evaluation via agents. Doc as below:
LLM Reasoning and Generation Benchmark. Evaluate LLMs in complex scenarios systematically.
[NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
[🏆Outstanding Paper Award at ACL 2024] MMToM-QA: Multimodal Theory of Mind Question Answering
(ACL 2025 Main) A Comprehensive Benchmark for Code Information Retrieval.
[ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
共 142 条 · 第 4 / 8 页