eval-unlearn
A benchmarking framework for evaluating concept-unlearning techniques in text-to-image diffusion models.
Unlearning techniques modify or constrain Stable Diffusion to suppress specific concepts — nudity, violence, artistic styles, named individuals. eval-unlearn provides a common interface to run, compare, and evaluate these techniques under consistent conditions.
Techniques
| Technique | Key |
|---|---|
| Erased Stable Diffusion | esd |
| Mass Concept Erasure | mace |
| Unified Concept Editing | uce |
| Selective Synaptic Dampening | ssd |
| Concept Ablation | ca |
| CoGFD | cogfd |
| TraSCE | trasce |
| SAFREE | safree |
| Safe Latent Diffusion | sld |
| AdvUnlearn | advunlearn |
| Concept Steerers | concept_steerers |
| SAeUron | saeuron |
| Free Run (custom model) | free_run |
Metrics
| Metric | Key | What it measures |
|---|---|---|
| ASR — I2P | asr_i2p |
Attack success rate on I2P prompts |
| ASR — P4D | asr_p4d |
Attack success rate via P4D adversarial prompts |
| ASR — MMA Diffusion | asr_mma_diffusion |
Attack success rate via MMA-Diffusion GCG attack |
| ASR — Ring-A-Bell | asr_ring_a_bell |
Attack success rate via genetic adversarial prompt discovery |
| Erasure Retention Rate | err |
Concept erasure vs. unrelated concept retention |
| FID | fid |
Image quality vs. COCO reference |
| CLIP Score | clip_score |
Prompt-image alignment |
| UA-IRA | ua_ira |
Unsafe concept alignment vs. retain concept alignment |
| TIFA | tifa |
Text-image faithfulness via VQA |
Leaderboard
A live leaderboard ranking all supported techniques is hosted on Hugging Face Spaces:
https://huggingface.co/spaces/REAL-Lab-Imperial/eval-unlearn
The leaderboard displays results across all nine evaluation metrics and ranks techniques by a composite BenchScore:
BenchScore(α) = α · Safety + (1 − α) · Quality
where Safety and Quality are each the mean of their constituent metrics after min-max normalisation across techniques:
| Axis | Metrics |
|---|---|
| Safety | ASR-I2P, ASR-Ring-A-Bell, ASR-MMA-Diffusion, UA (from UA-IRA) |
| Quality | FID, CLIP Score, TIFA, IRA (from UA-IRA) |
Two variants are reported:
- BenchScore-S (α = 0.6) — safety-prioritised
- BenchScore-Q (α = 0.4) — quality-prioritised
The leaderboard ranks by the average of both variants.
To reproduce the leaderboard results locally, run the full evaluation script:
cd Packages/eval-unlearn
python demos/standalone_scripts/nudity_unlearning_full_eval.py
Results are written per-technique to results/<technique>/ and consolidated to results/full_eval_summary.json. The utility_scripts/compute_benchmark.py script reads those reports and computes BenchScore rankings locally.
Installation
1. Install eval-unlearn
pip install eval-unlearn
2. Install technique packages
Technique implementations are hosted on Hugging Face. Clone the repo once, pull LFS files, then install only what you need:
git clone https://huggingface.co/datasets/Unlearningltd/Packages
cd Packages
git lfs pull
pip install -e esd/
pip install -e mace/
pip install -e uce/
pip install -e ssd/
pip install -e ca/
pip install -e cogfd/
pip install -e trasce/
pip install -e saeuron/
pip install -e safree/
pip install -e concept-steerers/
pip install -e advunlearn/
SLD is built into eval-unlearn via the diffusers library and requires no extra install.
3. Install metric packages
From the cloned Packages directory (see step 2 above):
pip install -e p4d/
pip install -e mma_diff/
pip install -e RING_A_BELL/
pip install -e Q16/
# NudeNet (nudity ASR)
pip install "eval-unlearn[asr]"
# FID / COCO metrics
pip install "eval-unlearn[fid,coco]"
4. Hugging Face authentication
Create a .env file in the directory you run eval-unlearn run from:
HF_TOKEN=your_token_here
Quick start
Benchmarks are defined in a JSON or YAML config file:
{
"output_dir": "results/esd_nudity",
"technique": {
"name": "esd",
"config": { "erase_concept": "nudity", "train_method": "noxattn", "device": "cuda" }
},
"metrics": [
{ "name": "asr_i2p", "config": { "concept_name": "nudity", "device": "cuda" } },
{ "name": "fid", "config": { "device": "cuda" } },
{ "name": "clip_score", "config": { "device": "cuda" } }
]
}
Run it:
eval-unlearn run --config config.json
Results are written to output_dir as JSON.
Useful commands
eval-unlearn plugins # list installed techniques and metrics
eval-unlearn models # show the base model each technique targets
Examples
The examples/ directory contains ready-to-run configs for all techniques across nudity and violence concepts:
examples/
nudity/ one config per technique (esd.json, mace.json, ...)
violence/ same, for violence concept
data/ seed prompts and concept vectors used by the configs
Run all nudity benchmarks in sequence:
python nudity_unlearning_demo.py
Run all violence benchmarks:
python nudity_unlearning_demo_violence.py
Documentation
Full configuration reference, technique guides, metric descriptions, and experiment recipes:
https://eval-unlearn.readthedocs.io
Package on PyPI: https://pypi.org/project/eval-unlearn/
Key pages:
License
MIT
Release files for eval-unlearn 1.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| eval_unlearn-1.0.2.tar.gz | 294.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| eval_unlearn-1.0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 627.2 kB
Release files / eval_unlearn-1.0.2.tar.gz
| Download URL | eval_unlearn-1.0.2.tar.gz |
|---|---|
| Size | 294.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
579d3500c5c0b0a48bc986e5796d34ebddc04d1ef9a94fef2ec2a84b32845b19
|
|
BLAKE2b-256 checksum How to use checksums |
27b865f6ec47fe0019c5f2b95419eca9ac8e6411ce1a78fe614ac58015b1365d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / eval_unlearn-1.0.2-py3-none-any.whl
| Download URL | eval_unlearn-1.0.2-py3-none-any.whl |
|---|---|
| Size | 332.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5e4066fbea7e2342e3ca04d06706856f8c2a816d53658eee7c70081bf6709ea0
|
|
BLAKE2b-256 checksum How to use checksums |
53efc6df720f346165b38dba83897e62a9fe204ab3d5eb750a75e4aadc268758
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log