benchmark
142 个项目 · ⭐ 197.5kCode for Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? [COLM 2024]
KAREN: Unifying Hatespeech Detection and Benchmarking
[ICML'26] Toward Human-like Audio-Visual Intelligence of Omni-MLLMs
H.E.I.M.D.A.L.L looks at fleet telemetry and gives you natural-language insights. GPU data loading (cuDF), local LLM inference (Gemma 2), and production NIM on GKE. Open the notebooks, run cells, get answers! Quick start should not take longer than 10 minutes and the T4 path is completely free!
An Interactive Game-based Vision Planning benchmark
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
This project implements 30+ variants of ANN algorithms to find the K nearest neighbors in high-dimensional vector spaces. It is meant as a convenient sandbox: drop in your own ANN code, run a one-liner, and instantly compare build/search speed and recall against the bundled baselines.
WMB-100K — The first 100,000-turn benchmark for AI memory systems
Pytorch code for "BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change"
EgoNormia | Benchmarking Physical Social Norm Understanding in VLMs
Benchmark OpenAI-compatible AI endpoints and AI Accelerators in a reproducible structured way
共 142 条 · 第 7 / 8 页