Evaluate unlearning beyond a single trained model
Repeating unlearning on one trained model measures variation conditional on that model. It does not reveal how the result changes when the original model is trained again. When training contributes to variability, extra unlearning runs cannot generally replace independent training runs. Lanyon et al. explain this through a training/unlearning variance decomposition and give guidance on allocating compute between the two.
SUPREME makes this experimental design practical: train I original models, run J unlearning repetitions per model, and optionally K evaluation repetitions per unlearned model. Separate stage seeds support investigating where variation arises. An extensible Python API supports comparisons against retraining, with execution from one GPU to a cluster.
Seed design and variance analysis · Research on training seeds
See what the paper reports
Existing Table 1 results on Pins Face Recognition, random-sample unlearning (forget 0.1%). Accuracy differences are unlearned minus retrained, in percentage points. Bars show one standard deviation across ten matched-seed pipelines (J = K = 1), combining variation across stages. These experiments used one GPU.
Interactive viewer · Download the tables · Paper · Presentation
Try the results example
Browse the published tables with Python 3.9 or later and its standard library:
git clone https://github.com/pedroandreou/supreme-unlearning.git
cd supreme-unlearning
python3 examples/paper_results.py
The results guide covers downloads, measurement definitions and exporting an offline viewer.
Framework comparison
SUPREME combines multi-seed image-unlearning evaluation with multi-GPU execution and configurable numerical precision.
| Framework | Domain in the comparison | Multi-seed | Multi-GPU | Multi-precision |
|---|---|---|---|---|
| OpenUnlearning | LLMs | Not shown | Yes | Yes |
| MUBox | Image classification | Not shown | Not shown | Not shown |
| ERASURE | Image classification | Yes | Not shown | Not shown |
| Deep Unlearn | Image classification | Yes | Not shown | Not shown |
| SUPREME | Image classification | Yes | Yes | Yes |
Feature definitions and sources.
🗃️ Available Components
| Component | Included |
|---|---|
| Datasets | CIFAR-10, CIFAR-20, CIFAR-100, Pins Face Recognition, Caltech-101 |
| Models | ResNet18, Vision Transformer |
| Methods | FT, Bad Teacher, Random Labels, UNSIR, SSD, LFSSD, SSD-Det, LFSSD-Det, ASSD, SCRUB, JIT |
| Reference models | Retrain and Original |
| Scenarios | Full-class, subclass and random-sample unlearning |
| Evaluation | Accuracy, membership inference, model distances and resource measurements |
Full component and hardware reference.
⚡ Quickstart
Install the Python library:
pip install supreme-unlearning
For a complete train → unlearn → evaluate run, follow the experiment quickstart. It covers environment setup, credentials and a small example.
The framework uses the paper's pinned dependency stack. Check the platform requirements and security guidance before loading external models or checkpoints.
📦 SUPREME as a Library
Register components from your own Python package:
import supreme
# Replace the module path with your implementation.
supreme.register_unlearning_method("mymethod", "your_package.your_method")
The library guide describes the public API and pipeline calls; the extension guide covers component interfaces.
📚 Documentation
| I want to… | Start here |
|---|---|
| Run local or SLURM experiments | Experiment guide |
| Reproduce the paper | Reproduction guide |
| Study variation across stages | Seed protocols |
| Add a method, metric, model or dataset | Custom-component notebook |
All documentation, including notation, implementation details, logging, tooling and maintainer workflows.
🤝 Contributing
Bug reports, new components and documentation contributions are welcome. Open an issue, read the contributing guide, or share a method and its results.
📝 Citing this work
If you use SUPREME in your research, please cite our paper and the original papers for the methods you use. See method credits and citations.
@misc{supreme2026,
title = {SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation},
author = {Petros Andreou, Jamie Lanyon, Axel Finke, Georgina Cosma},
year = {2026},
eprint = {2606.00380},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2606.00380}
}
This work was conducted at Loughborough University.
🙏 Acknowledgements
SUPREME builds on SSD, Bad Teacher and other open-source unlearning research. We thank their authors; full credits and citation guidance are available for each method.
📄 License
This project is licensed under the MIT License.
If SUPREME is useful for your research, star the repository to keep it handy.
Metadata
Release files for supreme-unlearning 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| supreme_unlearning-0.1.5.tar.gz | 162.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| supreme_unlearning-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 350.1 kB
Release files / supreme_unlearning-0.1.5.tar.gz
| Download URL | supreme_unlearning-0.1.5.tar.gz |
|---|---|
| Size | 162.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
366bd2ce21a142f1f9c36ee4ae3aeae355085a0ba7289bf4c50ceead5aad06ae
|
|
BLAKE2b-256 checksum How to use checksums |
91b667ec69da2e66da9915a7a00afad335309a80588e171fad1fd9a5fbcf5e55
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / supreme_unlearning-0.1.5-py3-none-any.whl
| Download URL | supreme_unlearning-0.1.5-py3-none-any.whl |
|---|---|
| Size | 187.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f0b41befb566a641d970219c9802675c24cdef6fd59d21463a69bb3b515cd0a1
|
|
BLAKE2b-256 checksum How to use checksums |
b97c9d240442697206bc01bd2dc2899555bb3fb1d8d61fa7b57daf5963116c7f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log