grpo
38 个项目 · ⭐ 71.7kA framework for agentic tool use training with reinforcement learning
Your AI intranet: network the computers you already own for inference and training.
NEWTON: Agentic Planning for Physically Grounded Video Generation
FlowSteer: agents designing agentic workflows via reinforced progressive canvas editing.
An implementation of GRPO for Unsloth's VLMs training
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.
🏆Winning Project | ModelGate is a contract-aware AI control plane that ingests customer contracts, extracts SLA/privacy/routing constraints, and generates an OpenAI-compatible endpoint that automatically routes every request to the optimal model. Simple queries go to cheap models. Complex queries go to premium ones.
[CVPR 2026 Highlight] ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
A lightweight post-training framework for LLMs and VLMs. 51 algorithms, 38 verified models. Scales with DeepSpeed, vLLM, and Ray.
Multi-node distributed LLM training framework
共 38 条 · 第 2 / 2 页