← 返回专题广场
reinforcement-learning
233 个项目 · ⭐ 1118.6k221
2025-01-20
最近推送
222
SCoRe: Training Language Models to Self-Correct via Reinforcement Learning
Python
⭐ 16
⑂ 0
· 2026-05-15推送
2026-05-15
最近推送
223
29 天前
最近推送
225
2025-02-02
最近推送
226
2026-04-22
最近推送
227
2024-01-31
最近推送
228
2026-05-22
最近推送
229
Reward a Language Model with pancakes 🥞
Jupyter Notebook
⭐ 11
⑂ 0
· 2023-09-29推送
2023-09-29
最近推送
230
2025-12-07
最近推送
231
2026-04-28
最近推送
232
Different post-training techniques for LLMs, including: SFT, DPO and Online RL
Python
⭐ 10
⑂ 3
MIT
· 2025-09-05推送
2025-09-05
最近推送
233
2025-04-28
最近推送
共 233 条 · 第 12 / 12 页