evaluation
101 个项目 · ⭐ 261.3k[ACL 2024] User-friendly evaluation framework: Eval Suite & Benchmarks: UHGEval, HaluEval, HalluQA, etc.
This repo is deprecated. Please go to langchain-ai/docs.
Dialectical reasoning architecture for LLMs (Thesis → Antithesis → Synthesis)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Automated system for LLM evaluation via agents. Doc as below:
[ICLR'26, NAACL'25 Demo] Toolkit & Benchmark for evaluating the trustworthiness of generative foundation models.
Language AI Engineering Lab, a place where you can deeply understand and build modern Language AI systems, from fundamentals to production.
Rank LLMs, RAG systems, and prompts using automated head-to-head evaluation
⚡ Unit tests for AI. Test prompts, compare models, save money.
This repo contains the code for "MEGA-Bench Scaling Multimodal Evaluation to over 500 Real-World Tasks" [ICLR 2025]
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Adaptive Reasoning Engine for Efficient and Context-Aware Intelligence
共 101 条 · 第 4 / 6 页