← 返回专题广场
mechanistic-interpretability
10 个项目 · ⭐ 8871
To know what models don't say out loud.
HTML
⭐ 238
⑂ 21
NOASSERTION
· 2026-07-23推送
2026-07-23
最近推送
2
Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.
Python
⭐ 224
⑂ 22
Apache-2.0
· 21 天前推送
21 天前
最近推送
3
Steering vectors for transformer language models in Pytorch / Huggingface
Python
⭐ 162
⑂ 19
MIT
· 2025-02-22推送
2025-02-22
最近推送
4
Generating and validating natural-language explanations for the brain.
Jupyter Notebook
⭐ 66
⑂ 12
MIT
· 2026-06-02推送
2026-06-02
最近推送
5
Sparse and discrete interpretability tool for neural networks
Python
⭐ 64
⑂ 5
MIT
· 2024-02-12推送
2024-02-12
最近推送
6
15 天前
最近推送
7
2026-07-16
最近推送
8
Lightweight representation engineering dataflow operations for agent developers.
Python
⭐ 23
⑂ 1
Apache-2.0
· 2026-05-28推送
2026-05-28
最近推送
9
2026-07-15
最近推送