← 返回专题广场
reinforcement-learning-from-human-feedback
3 个项目 · ⭐ 11.9k1
9 天前
最近推送
2
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Python
⭐ 1.6k
⑂ 133
Apache-2.0
· 2025-11-24推送
2025-11-24
最近推送
3
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
Python
⭐ 336
⑂ 22
Apache-2.0
· 2025-04-24推送
2025-04-24
最近推送