← 返回专题广场
post-training
25 个项目 · ⭐ 7.0k1
2026-06-19
最近推送
2
2026-07-06
最近推送
3
2026-07-10
最近推送
5
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Python
⭐ 578
⑂ 141
Apache-2.0
· 2 天前推送
2 天前
最近推送
6
2026-04-25
最近推送
7
18 天前
最近推送
8
Train a Language Model with GRPO to create a schedule from a list of events and priorities
Jupyter Notebook
⭐ 273
⑂ 15
Apache-2.0
· 2026-04-08推送
2026-04-08
最近推送
9
2026-04-30
最近推送
10
RapidFire AI: Rapid AI Customization from RAG to Fine-Tuning
JavaScript
⭐ 167
⑂ 27
Apache-2.0
· 3 天前推送
3 天前
最近推送
11
17 天前
最近推送
12
[ECCV 2026🔥] SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
Python
⭐ 93
⑂ 7
Apache-2.0
· 2025-11-26推送
2025-11-26
最近推送
13
A High-Efficiency System of Large Language Model Based Search Agents
Python
⭐ 80
⑂ 5
· 2025-07-02推送
2025-07-02
最近推送
14
2025-11-24
最近推送
15
2026-02-01
最近推送
16
2025-04-01
最近推送
17
[EMNLP 2022] Continual Training of Language Models for Few-Shot Learning
Python
⭐ 44
⑂ 1
· 2023-02-13推送
2023-02-13
最近推送
18
2025-03-16
最近推送
19
[AAAI 2026] D²PPO: Diffusion Policy Policy Optimization with Dispersive Loss.
Python
⭐ 42
⑂ 4
MIT
· 2025-11-22推送
2025-11-22
最近推送
20
Code repository for "Post-pre-training for Modality Alignment in Vision-Language Foundation Models" (CVPR2025)
Python
⭐ 41
⑂ 2
NOASSERTION
· 2025-07-25推送
2025-07-25
最近推送
共 25 条 · 第 1 / 2 页