d-f

d-f/llm-summarization

LoRA supervised fine-tuning, RLHF (PPO) and RAG with llama-3-8B on the TLDR summarization dataset

⭐ 14 ⑂ 1 Python · 2025-02-02推送
14
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
2025-02-02
最近推送
原文 中文
暂无 README