benchmark
142 个项目 · ⭐ 197.5kPhyX: Does Your Model Have the "Wits" for Physical Reasoning?
:hugs: AeroPath: An airway segmentation benchmark dataset with challenging pathology
Generate degraded speech datasets for noise-robust ASR benchmarking
This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"
ICLR 2026 | ChemEval: 4-level, 13-dimension, 62-task text/multimodal chemistry benchmark for evaluating LLMs and MLLMs.
Lynn GitHub 镜像仓 · Primary repository: https://github.com/MerkyorLynn/Lynn · Downloads: https://download.merkyorlynn.com/download.html
A lightweight library for normalizing speech transcripts before computing WER
For our ACL25 Paper: Can Language Models Replace Programmers? RepoCod Says ‘Not Yet’ - by Shanchao Liang and Yiran Hu and Nan Jiang and Lin Tan
Probing the limitations of multimodal language models for chemistry and materials research
共 142 条 · 第 6 / 8 页