adversarial-attacks
19 个项目 · ⭐ 59.1kA Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX
A unified evaluation framework for large language models
PyTorch implementation of adversarial attacks [torchattacks]
Raising the Cost of Malicious AI-Powered Image Editing
Anti-DreamBooth: Protecting users from personalized text-to-image synthesis (ICCV 2023)
[ICLR 2025] Dissecting adversarial robustness of multimodal language model agents
The official implementation of ECCV'24 paper "To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now". This work introduces one fast and effective attack method to evaluate the harmful-content generation ability of safety-driven unlearned diffusion models.
An ASR (Automatic Speech Recognition) adversarial attack repository.
Sparse Autoencoders (SAE) vs CLIP fine-tuning fun.