d-f/llm-summarization
LoRA supervised fine-tuning, RLHF (PPO) and RAG with llama-3-8B on the TLDR summarization dataset
⭐ 14
⑂ 1
Python
· 2025-02-02推送
14
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
2025-02-02
最近推送
原文
中文
暂无 README