← 返回专题广场
ai-safety
54 个项目 · ⭐ 33.0k41
Benchmarking Open-Ended Inference Optimization by AI Agents
Python
⭐ 41
⑂ 6
Apache-2.0
· 2026-07-07推送
2026-07-07
最近推送
42
2026-04-30
最近推送
43
2026-07-17
最近推送
44
一份来自2026年的AI精神病理学诊断报告 / A Pathological Diagnosis of AI Civilization
⭐ 37
⑂ 0
NOASSERTION
· 2026-01-23推送
2026-01-23
最近推送
45
13 天前
最近推送
46
Stop LLMs from hallucinating your guesses as facts. Clarity Gate is a verification protocol for your documents that are going to be provided to LLMs or RAG systems. Place automatically the missing uncertainty markers to avoid confident hallucinations. HITL for non-directly verifiable claims.
Python
⭐ 32
⑂ 3
NOASSERTION
· 2026-03-02推送
2026-03-02
最近推送
47
2025-06-11
最近推送
48
4 小时前
最近推送
49
Lightweight, Self-Hosted AI Guardrails Model based on ModernBERT.
Jupyter Notebook
⭐ 12
⑂ 1
Apache-2.0
· 2026-03-25推送
2026-03-25
最近推送
50
Runtime detector for reward hacking and misalignment in LLM agents (89.7% F1 on 5,391 trajectories).
Python
⭐ 12
⑂ 1
Apache-2.0
· 26 天前推送
26 天前
最近推送
51
[ICML 2025] Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
Python
⭐ 12
⑂ 0
· 2025-06-11推送
2025-06-11
最近推送
52
13 天前
最近推送
53
2026-06-30
最近推送
54
2026-07-15
最近推送
共 54 条 · 第 3 / 3 页