← 返回专题广场
harmful
6 个项目 · ⭐ 4371
2026-06-22
最近推送
2
This is the official code for the paper "Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation"
Python
⭐ 56
⑂ 4
Apache-2.0
· 2025-02-02推送
2025-02-02
最近推送
3
This is the official code for the paper "Vaccine: Perturbation-aware Alignment for Large Language Models" (NeurIPS2024)
Shell
⭐ 52
⑂ 6
Apache-2.0
· 2026-01-15推送
2026-01-15
最近推送
4
2025-03-23
最近推送
5
This is the official code for the paper "Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning" (NeurIPS2024)
Python
⭐ 29
⑂ 0
Apache-2.0
· 2024-09-11推送
2024-09-11
最近推送
6
2025-07-15
最近推送