← 返回专题广场
safety
14 个项目 · ⭐ 21.0k1
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Python
⭐ 7.0k
⑂ 812
NOASSERTION
· 1 天前推送
1 天前
最近推送
2
4 小时前
最近推送
3
2026-06-17
最近推送
4
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Python
⭐ 1.6k
⑂ 133
Apache-2.0
· 2025-11-24推送
2025-11-24
最近推送
5
12 天前
最近推送
6
Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。
⭐ 1.2k
⑂ 89
Apache-2.0
· 2024-02-27推送
2024-02-27
最近推送
7
2026-06-05
最近推送
8
AutoHarness: Automated Harness Engineering for AI Agents
Python
⭐ 368
⑂ 28
MIT
· 2026-04-03推送
2026-04-03
最近推送
9
2026-06-22
最近推送
10
[ICLR 2025] Dissecting adversarial robustness of multimodal language model agents
Python
⭐ 140
⑂ 10
MIT
· 2025-02-20推送
2025-02-20
最近推送
11
This is the official code for the paper "Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation"
Python
⭐ 56
⑂ 4
Apache-2.0
· 2025-02-02推送
2025-02-02
最近推送
12
2025-03-23
最近推送
13
[WSDM 2026] LookAhead Tuning: Safer Language Models via Partial Answer Previews
Python
⭐ 17
⑂ 0
MIT
· 2025-12-14推送
2025-12-14
最近推送
14
2025-07-15
最近推送