SuleynanAuir/Distributed-Fine-Tuning-of-Qwen2.5-7B-as-HotelService-Master
Hotel-domain conversational assistant using Qwen2.5-7B-Instruct, focusing on practical, reproducible fine-tuning. It supports SFT + LoRA for a stable baseline, SFT + QLoRA for 7B model training with 4-bit quantization under limited GPU memory, and DPO to enhance responses via human preference pairs
⭐ 11
⑂ 0
Python
MIT
· 2026-03-09推送
11
Watchers
0
贡献者
0
Commits
0
Releases
0
Open Issues
2026-03-09
最近推送
原文
中文
暂无 README