rlhf
60 个项目 · ⭐ 204.7kRecipes to train reward model for RLHF.
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
Multimodal RL training framework for diffusion & omni models
聚宝盆(Cornucopia): 中文金融系列开源可商用大模型,并提供一套高效轻量化的垂直领域LLM训练框架(Pretraining、SFT、RLHF、Quantize等)
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
The official implementation of Self-Play Preference Optimization (SPPO)
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
MindSpore online courses: Step into LLM
The Official Repo for "Quick Start Guide to Large Language Models"
🛰️ 基于真实医疗对话数据在ChatGLM上进行LoRA、P-Tuning V2、Freeze、RLHF等微调,我们的眼光不止于医疗问答
Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Part of the Gaslamp AI platform.
Python client library for improving your LLM app accuracy
共 60 条 · 第 2 / 3 页