Skip to main content

Evaluate unlearning beyond a single trained model

Repeating unlearning on one trained model measures variation conditional on that model. It does not reveal how the result changes when the original model is trained again. When training contributes to variability, extra unlearning runs cannot generally replace independent training runs. Lanyon et al. explain this through a training/unlearning variance decomposition and give guidance on allocating compute between the two.

SUPREME makes this experimental design practical: train I original models, run J unlearning repetitions per model, and optionally K evaluation repetitions per unlearned model. Separate stage seeds support investigating where variation arises. An extensible Python API supports comparisons against retraining, with execution from one GPU to a cluster.

Seed design and variance analysis · Research on training seeds

See what the paper reports

Published forget-accuracy differences, showing means and standard deviations across ten seeds

Existing Table 1 results on Pins Face Recognition, random-sample unlearning (forget 0.1%). Accuracy differences are unlearned minus retrained, in percentage points. Bars show one standard deviation across ten matched-seed pipelines (J = K = 1), combining variation across stages. These experiments used one GPU.

Interactive viewer · Download the tables · Paper · Presentation

Try the results example

Browse the published tables with Python 3.9 or later and its standard library:

git clone https://github.com/pedroandreou/supreme-unlearning.git
cd supreme-unlearning
python3 examples/paper_results.py

The results guide covers downloads, measurement definitions and exporting an offline viewer.

Framework comparison

SUPREME combines multi-seed image-unlearning evaluation with multi-GPU execution and configurable numerical precision.

Framework Domain in the comparison Multi-seed Multi-GPU Multi-precision
OpenUnlearning LLMs Not shown Yes Yes
MUBox Image classification Not shown Not shown Not shown
ERASURE Image classification Yes Not shown Not shown
Deep Unlearn Image classification Yes Not shown Not shown
SUPREME Image classification Yes Yes Yes

Feature definitions and sources.

🗃️ Available Components

Component Included
Datasets CIFAR-10, CIFAR-20, CIFAR-100, Pins Face Recognition, Caltech-101
Models ResNet18, Vision Transformer
Methods FT, Bad Teacher, Random Labels, UNSIR, SSD, LFSSD, SSD-Det, LFSSD-Det, ASSD, SCRUB, JIT
Reference models Retrain and Original
Scenarios Full-class, subclass and random-sample unlearning
Evaluation Accuracy, membership inference, model distances and resource measurements

Full component and hardware reference.

⚡ Quickstart

Install the Python library:

pip install supreme-unlearning

For a complete train → unlearn → evaluate run, follow the experiment quickstart. It covers environment setup, credentials and a small example.

The framework uses the paper's pinned dependency stack. Check the platform requirements and security guidance before loading external models or checkpoints.

📦 SUPREME as a Library

Register components from your own Python package:

import supreme

# Replace the module path with your implementation.
supreme.register_unlearning_method("mymethod", "your_package.your_method")

The library guide describes the public API and pipeline calls; the extension guide covers component interfaces.

📚 Documentation

I want to… Start here
Run local or SLURM experiments Experiment guide
Reproduce the paper Reproduction guide
Study variation across stages Seed protocols
Add a method, metric, model or dataset Custom-component notebook

All documentation, including notation, implementation details, logging, tooling and maintainer workflows.

🤝 Contributing

Bug reports, new components and documentation contributions are welcome. Open an issue, read the contributing guide, or share a method and its results.

📝 Citing this work

If you use SUPREME in your research, please cite our paper and the original papers for the methods you use. See method credits and citations.

@misc{supreme2026,
  title  = {SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation},
  author = {Petros Andreou, Jamie Lanyon, Axel Finke, Georgina Cosma},
  year   = {2026},
  eprint = {2606.00380},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url    = {https://arxiv.org/abs/2606.00380}
}

This work was conducted at Loughborough University.

🙏 Acknowledgements

SUPREME builds on SSD, Bad Teacher and other open-source unlearning research. We thank their authors; full credits and citation guidance are available for each method.

📄 License

This project is licensed under the MIT License.

If SUPREME is useful for your research, star the repository to keep it handy.

Metadata

Release files for supreme-unlearning 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for supreme-unlearning 0.1.5
File Size Uploaded
supreme_unlearning-0.1.5.tar.gz 162.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for supreme-unlearning 0.1.5
File Interpreter ABI Platform
supreme_unlearning-0.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 350.1 kB

Release files / supreme_unlearning-0.1.5.tar.gz

Download URL supreme_unlearning-0.1.5.tar.gz
Size 162.9 kB
Tags Source
SHA-256 checksum
How to use checksums
366bd2ce21a142f1f9c36ee4ae3aeae355085a0ba7289bf4c50ceead5aad06ae
BLAKE2b-256 checksum
How to use checksums
91b667ec69da2e66da9915a7a00afad335309a80588e171fad1fd9a5fbcf5e55
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / supreme_unlearning-0.1.5-py3-none-any.whl

Download URL supreme_unlearning-0.1.5-py3-none-any.whl
Size 187.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f0b41befb566a641d970219c9802675c24cdef6fd59d21463a69bb3b515cd0a1
BLAKE2b-256 checksum
How to use checksums
b97c9d240442697206bc01bd2dc2899555bb3fb1d8d61fa7b57daf5963116c7f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.5 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page