evaluation
101 个项目 · ⭐ 261.3kEvaluate results from ASR/Speech-to-Text quickly
This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"
Latxa: An Open Language Model and Evaluation Suite for Basque
A starter kit for evaluating benchmarks on the 🤗 Hub
[ICLR 2026🔥] SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]
⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.
GenAI - Building AI Workflows | Agentic Appls | Working with LLMs | HuggingFace | OpenAI | Gemini
WMB-100K — The first 100,000-turn benchmark for AI memory systems
A comprehensive framework to explore whether embodied multimodal models are plausibly resilient
共 101 条 · 第 5 / 6 页