← 返回专题广场
dpo
24 个项目 · ⭐ 29.0k21
Hotel-domain conversational assistant using Qwen2.5-7B-Instruct, focusing on practical, reproducible fine-tuning. It supports SFT + LoRA for a stable baseline, SFT + QLoRA for 7B model training with 4-bit quantization under limited GPU memory, and DPO to enhance responses via human preference pairs
Python
⭐ 11
⑂ 0
MIT
· 2026-03-09推送
2026-03-09
最近推送
22
16 天前
最近推送
23
Different post-training techniques for LLMs, including: SFT, DPO and Online RL
Python
⭐ 10
⑂ 3
MIT
· 2025-09-05推送
2025-09-05
最近推送
共 24 条 · 第 2 / 2 页