Skip to main content

Exhibit: Command line tool to create anonymised demonstrator data

The goal of Exhibit is to make it easier to generate synthethic data at scale in a controlled and reproducible way.


main

Build Status CodeQL codecov

latest

Build Status CodeQL codecov

Key features

  • Control all aspects of the anonymisation process: which columns to anonymise and to what degree
  • Rapidly iterate on the anonymisation options
  • Set categorical weights to create custom distributions
  • Use regular expressions to bulk-anonymise identifiers
  • Add columns derived from newly anonymised data
  • Preserve important relationships between your columns (paired, hierarchical, custom)
  • Add outliers to any subset of the generated data
  • Generate and manipulate missing data and timeseries
  • Generate geo-spatial data using H3 hexes
  • Augment your synthetic data with compiled machine learning models and custom functions
  • Use SQL to generate conditional values based on external tables

Installation:

To install using pip, enter the following command at a Bash or Windows command prompt:

pip install exhibit

Alternatively, download or clone the repository and run pip install . from the root folder.

Quickstart

Exhibit has two principal modes of operation:

  • fromdata produces a detailed, user-editable .yml specification
  • fromspec which produces the anonymised dataset from the supplied specification

See the -h listing for the full list of optional command line parameters.

The repository includes a few sample datasets and specifications.
You can find them in exhibit/sample/_data and exhibit/sample/_spec

To create a demo dataset, run:
exhibit fromspec exhibit/sample/_spec/inpatients_demo.yml -o demo.csv

To create a demo specification that equialises all probabilities and weights, run:
exhibit fromdata exhibit/sample/_data/inpatients.csv -ew -o demo.yml

Database

Exhibit is bundled with a SQLite3 database and a Python utility tool to interact with it. Alternatively, you can connect directly to /exhbit/db/exhibit.db. The database contains three sample aliasing datasets: mountains, birds and patients designed to help you quickly alias original values without manually editing individual column values.

  • mountains has 15 mountain ranges and their top 10 peaks making it useful for aliasing hierarchical pairs, like NHS Boards and Hospitals.
  • birds has 150 pairs of common / scientific bird names. This can be useful for 1:1 paired columns.
  • patients has 360 made-up patient records with details such as gender, 5-year age band, date of birth and CHI number. Fields from this dataset can be selectively pulled in when linked data is required.
  • dates has dates ranging from 1900-01-01 to 2100-01-01 at a single day interval. This table is useful if you have a SQL statement in the anonymising_set that picks dates based on a condition.

The database is also used to store temporary data for columns where the number of unique values exceeds user threshold and thus not available for editing directly in the yml file.

Note that original, confidential data might be saved in the exhibit/db/exhibit.db file on your local machine. You can purge all temporary tables by calling --purge command from the included utility tool or by interfacing with the database directly.

Disclaimer

Please note that the degree of anonymisation for each dataset produced by the tool will depend heavily on user choices in the specification. As such, there is no guarantee that confidential data will be suitably masked under all scenarios. If you intend to work with sensitive data, make sure to thoroughly evaluate the output before making it public.

Metadata

Release files for exhibit 0.9.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for exhibit 0.9.9
File Size Uploaded
exhibit-0.9.9.tar.gz 870.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for exhibit 0.9.9
File Interpreter ABI Platform
exhibit-0.9.9-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / exhibit-0.9.9.tar.gz

Download URL exhibit-0.9.9.tar.gz
Size 870.0 kB
Tags Source
SHA-256 checksum
How to use checksums
6d0ca635ae1a9e7e5a9aabdc887b749e6a5e530fd248adeba9ebb4b69e20bd05
BLAKE2b-256 checksum
How to use checksums
5b3881f5bc1c83c469f48a3daaaffd814be10a0ba5a6cb38870b9acaaa8ba9fc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.11

Release files / exhibit-0.9.9-py3-none-any.whl

Download URL exhibit-0.9.9-py3-none-any.whl
Size 904.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
245708e89782853cbc55d61a3b808fafb772cc5e3f63851c8763504e934e98ad
BLAKE2b-256 checksum
How to use checksums
e1a821c5f0d8edf20aefc05fa4564160d253af51ba7c81b37c7afb76dfb1da30
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.11

Release history Release notifications | RSS feed

This release

0.9.9 This release

2 release files

0.9.8

2 release files

0.9.7

2 release files

0.9.6

2 release files

0.9.5

2 release files

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page