rag-evaluation
12 个项目 · ⭐ 15.0k🐢 Open-Source Evaluation & Testing library for LLM Agents
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
Dataset and benchmark for RAG on company internal documents.
中文优先的企业 RAG 知识库:可控解析、治理、切块、混合检索、重排、引用、图谱、评测与 Dify 接入。
RAG evaluation without the need for "golden answers"
Red Teaming python-framework for testing chatbots and GenAI systems.
RAG boilerplate with semantic/propositional chunking, hybrid search (BM25 + dense), LLM reranking, query enhancement agents, CrewAI orchestration, Qdrant vector search, Redis/Mongo sessioning, Celery ingestion pipeline, Gradio UI, and an evaluation suite (Hit-Rate, MRR, hybrid configs).
⚡️ The "1-Minute RAG Audit" — Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, async, visual reports.