benchmark
142 个项目 · ⭐ 197.5kA 13B large language model developed by Baichuan Intelligent Technology
A unified evaluation framework for large language models
CPU-X is a Free software that gathers information on CPU, motherboard and more
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
Efficient Retrieval Augmentation and Generation Framework
[CVPR2024 Highlight] VBench - We Evaluate Video Generation
Quickly find bottlenecks in Rust - one profiler for CPU, memory, SQL, HTTP, I/O and async code.
ADR secures enterprise AI agents through observability, security benchmarking, and threat detection. Deployed at Uber.
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
共 142 条 · 第 2 / 8 页