grpo
38 个项目 · ⭐ 71.7kSolve Visual Understanding with Reinforced VLMs
Implement a reasoning LLM in PyTorch from scratch, step by step
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Multimodal RL training framework for diffusion & omni models
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Agentic RAG R1 Framework via Reinforcement Learning
Codebase of GRPO: Implementations and Resources of GRPO and Its Variants
Zero-friction LLM fine-tuning skill for Claude Code, Gemini CLI & any ACP agent. Unsloth on NVIDIA · TRL+MPS/MLX on Apple Silicon. Automates env setup, LoRA training (SFT, DPO, GRPO, vision), post-hoc GRPO log diagnostics, evaluation, and export end-to-end. Part of the Gaslamp AI platform.
Train a Language Model with GRPO to create a schedule from a list of events and priorities
Official implementation of GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
共 38 条 · 第 1 / 2 页