DP+Fair Benchmarking Framework
This repository provides a Python framework for benchmarking fairness mechanisms on Differentially Private Synthetic Data.
Features
- ⚡ Simple, reproducible setup for benchmarking algorithms
- 🧩 Flexible API to plug in any classifier implementing
fit,predict, andpredict_proba - 📊 Pre-offered datasets included under
data/ - 🔬 Configurable experiment settings: dataset schema, dataset synthesizer, seeds, privacy-budget, input/outputs, classifier, data pre-processing.
Installation
First, we strongly recommend creating a new Python environment before installing the package. This helps maintain a clean and reproducible setup, facilitates dependency and version management and update, and minimises potential conflicts with previously installed libraries.
python3 -m venv dpfair-env
source dpfair-env/bin/activate # On Windows: dpfair-env\\Scripts\\activate
# Installation approach via PyPi or source installation
When using Python 3.9, a specific version of the MBI library (private-pgm) must be installed by running the following command:
pip install git+https://github.com/ryan112358/private-pgm.git@01f02f17eba440f4e76c1d06fa5ee9eed0bd2bca
If you are a Windows or Linux user, please install PyTorch CPU-only dependencies: [Step not required for Mac users]
pip install torch --index-url https://download.pytorch.org/whl/cpu
Users on MacOS are required to install llvm-openmp, either using conda or brew, as below:
Conda:
conda install -c conda-forge llvm-openmp
Brew:
brew install libomp
Then, the library installation can proceed as per usual.
Our recommended method is via PyPI. With a Python 3.9+ and <3.13 environment:
pip install BenchmarkDPFair
Alternatively, there is the possibility to install from source, which also gives access to the data/, example/, and notebook/ directories (be aware that these may consume more resources on your disk given the data provided):
git clone https://github.com/vinicius-verona/dp-fair-intervention-benchmark.git
cd dp-fair-intervention-benchmark
pip install -e .
Repository Structure
├── data/ # Pre-offered datasets
├── src/ # Core source code
├── example/ # Files used in the main paper and a dummy example.
└── README.md
Quick Start
Here is a dummy example:
import argparse
from typing import List, Union
from BenchmarkDPFair.DataGenerator import generate_data, DatasetGeneratorConfig
from BenchmarkDPFair.Benchmark import benchmark, BenchmarkInfo, BenchmarkDatasetConfig
from sklearn.linear_model import LogisticRegression
ESTIMATOR_PARAMS = {
'max_iter': 10000,
'solver': 'saga',
'l1_ratio': 0.5,
'C': 0.8
}
lr = LogisticRegression
classifiers = [lr]
ckwargs = [
ESTIMATOR_PARAMS,
]
classifier_name = ["LR"]
combinations = [
(0, 0),
(0, 1),
]
synths = ["aim", "mst"]
if __name__ == "__main__":
parser = argparse.ArgumentParser(description="Arguments of Data Generation for Adult")
parser.add_argument(
"--seeds", "-s",
nargs="+", # 1 or more values
type=int # convert automatically to int
)
args = parser.parse_args()
seeds = args.seeds
eps : List[Union[int,float]] = [.05, .1]
for synthesizer in synths:
for s in seeds:
data_conf = DatasetGeneratorConfig(
name = "Compas",
target= "two_year_recid",
synthesizer = synthesizer,
root_dir="./data",
sensitive_attr = "race",
categorical_cols = ['race', 'score_text', 'c_charge_degree','age', 'sex', 'two_year_recid'],
sensitive_cols = ['race', 'sex'],
ordinal_cols = ['priors_count'],
privacy_budgets=eps,
binary_encoder=binary_encode,
seed = s,
test_split_size=0.4,
data_filter = filter_compas
)
generate_data(f"compas.csv", "", data_conf, "./data", verbose=True)
for clf_idx, syn_idx in combinations:
classifier = classifiers[clf_idx]
synth = synths[syn_idx]
benchmark_config = BenchmarkInfo(
dp_method=synth,
output_dir=f"./output/Dummy-Compas/{classifier_name[clf_idx]}/",
seeds=seeds,
eps = eps,
classifier=classifier,
classifier_kwargs=ckwargs[clf_idx]
)
benchmark_dataset = BenchmarkDatasetConfig(
name = "Compas",
target= "two_year_recid",
root_dir="./data",
sensitive_attr = "race",
index_col="Unnamed: 0",
categorical_cols = ['race', 'score_text', 'c_charge_degree','age', 'sex', 'two_year_recid'],
ordinal_cols=["priors_count"],
sensitive_cols = ['race', 'sex'],
)
benchmark(benchmark_info=benchmark_config, data_conf=benchmark_dataset)
More detailed examples can be found in the example/ directory.
License
License: MIT
Contributing
Contributions are welcome:
- Open an issue for bug reports or feature requests
Metadata
Release files for BenchmarkDPFair 0.2.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| benchmarkdpfair-0.2.7.tar.gz | 24.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| benchmarkdpfair-0.2.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 52.0 kB
Release files / benchmarkdpfair-0.2.7.tar.gz
| Download URL | benchmarkdpfair-0.2.7.tar.gz |
|---|---|
| Size | 24.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c94d9b08065c8b651ba197dee6d6c1ca23acc99d74548e425505c6a1067fc2a5
|
|
BLAKE2b-256 checksum How to use checksums |
e74c4d8ae7de2e2758acca60af9d0c6a3c907a219d8f4bb8650c0ac86e1ba25d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / benchmarkdpfair-0.2.7-py3-none-any.whl
| Download URL | benchmarkdpfair-0.2.7-py3-none-any.whl |
|---|---|
| Size | 28.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5ac62389f7405c351833fc1fa1241464eb601436a54e26080dd402bb1f2c9c51
|
|
BLAKE2b-256 checksum How to use checksums |
ab422d3cc8c977a67b64e92ca88fe6eaa221bbbe3b90e388d098525679e3c0ac
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|