interpretability
51 个项目 · ⭐ 102.2kChat2Graph: Graph Native Agentic System.
The Truth Is In There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
Diffusers-Interpret 🤗🧨🕵️♀️: Model explainability for 🤗 Diffusers. Get explanations for your generated images.
Interpretable Causal Diffusion Language Models
To know what models don't say out loud.
Jacobian-Brainwash : A manual alignment tool for large language models built on Anthropic's Jacobian Lens. Results are exportable.
A python package for benchmarking interpretability techniques on Transformers.
Interpret text data with LLMs (sklearn compatible).
A library for finding knowledge neurons in pretrained transformer models.
Collection of NLP model explanations and accompanying analysis tools
[ICLR 2026] DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
Robust multimodal image registration via keypoints
Generating and validating natural-language explanations for the brain.
Sparse and discrete interpretability tool for neural networks
An interpretable foundation model of the patient clinical timeline: event forecasting, calibrated time-to-event alerts, and concept-level interpretability, benchmarked head-to-head against tuned GBMs, tabular foundation models, and survival baselines on MIMIC-IV, eICU, and GEMINI via the MEDS standard.
Time series explainability via self-supervised model behavior consistency
Code for the paper "Aligning LLM Agents by Learning Latent Preference from User Edits".
Interpretable and efficient predictors using pre-trained language models. Scikit-learn compatible.
共 51 条 · 第 2 / 3 页