rlhf
60 个项目 · ⭐ 204.7kUnified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
The official GitHub page for the survey paper "A Survey of Large Language Models".
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
Robust recipes to align language models with human and AI preferences
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Implement a reasoning LLM in PyTorch from scratch, step by step
Align Anything: Training All-modality Model with Feedback
A curated list of reinforcement learning with human feedback resources (continually updated)
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)
共 60 条 · 第 1 / 3 页