Skip to main content

Clean personally identifiable information from dirty dirty text.

Project description

Remove personally identifiable information from free text. Sometimes we have additional metadata about the people we wish to anonymize. Other times we don’t. This package makes it easy to seamlessly scrub personal information from free text, without compromising the privacy of the people we are trying to protect.

scrubadub currently supports removing:

  • Names

  • Email addresses

  • Addresses/Postal codes (US, GB, CA)

  • Credit card numbers

  • Dates of birth

  • URLs

  • Phone numbers

  • Username and password combinations

  • Skype/twitter usernames

  • Social security numbers (US and GB national insurance numbers)

  • Tax numbers (GB)

  • Driving licence numbers (GB)

Build Status Version Downloads Test Coverage Documentation Status

Quick start

Getting started with scrubadub is as easy as pip install scrubadub and incorporating it into your python scripts like this:

>>> import scrubadub

# My cat may be more tech-savvy than most, but he doesn't want other people to know it.
>>> text = "My cat can be contacted on example@example.com, or 1800 555-5555"

# Replaces the phone number and email addresse with anonymous IDs.
>>> scrubadub.clean(text)
'My cat can be contacted on {{EMAIL}}, or {{PHONE}}'

There are many ways to tailor the behavior of scrubadub using different Detectors and PostProcessors. Scrubadub is highly configurable and supports localisation for different languages and regions.

Installation

To install scrubadub using pip, simply type:

pip install scrubadub

There are several other packages that can optionally be installed to enable extra detectors. These scrubadub_address, scrubadub_spacy and scrubadub_stanford, see the relevant documentation (address detector documentation and name detector documentation) for more info on these as they require additional dependencies. This package requires at least python 3.6. For python 2.7 or 3.5 support use v1.2.2 which is the last version with support for these versions.

New maintainers

LeapBeyond are excited to be supporting scrubadub with ongoing maintenance and development. Thanks to all of the contributors who made this package a success, but especially @deanmalmgren, IDEO and Datascope.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrubadub-2.0.1.tar.gz (46.6 kB view details)

Uploaded Source

Built Distribution

scrubadub-2.0.1-py3-none-any.whl (65.2 kB view details)

Uploaded Python 3

File details

Details for the file scrubadub-2.0.1.tar.gz.

File metadata

  • Download URL: scrubadub-2.0.1.tar.gz
  • Upload date:
  • Size: 46.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.2 CPython/3.11.4

File hashes

Hashes for scrubadub-2.0.1.tar.gz
Algorithm Hash digest
SHA256 52a1fb8aa9bc0226043e02c3ec22d450bd4ebeede9e7e8db2def7c89b37c5aad
MD5 67c901dc153479682a39a3f415b97275
BLAKE2b-256 6f24f56c1b27689eff1809791b37660a9b1687ddfb157c0e380114245d67af1b

See more details on using hashes here.

File details

Details for the file scrubadub-2.0.1-py3-none-any.whl.

File metadata

  • Download URL: scrubadub-2.0.1-py3-none-any.whl
  • Upload date:
  • Size: 65.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.2 CPython/3.11.4

File hashes

Hashes for scrubadub-2.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 44b9004998a03aff4c6b5d9073a52895081742f994470083a7be610b373e62b7
MD5 4642d4a1cb79d070134f516d93a0efba
BLAKE2b-256 f4c504b959566c85914b17327e40d25b0535b0209a5a5216006443b769bebe25

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page