datapatch
A Python library for defining rule-based overrides on messy data. Imagine, for example,
trying to import a dataset in each row is associated with a country - which have been
entered by humans. You might find country names like Northkorea, or Greet Britain
that you want to normalise. datapatch creates a mechanism to build a flexible lookup
table (usually stored as a YAML file) to catch and repair these data issues.
Installation
You can install datapatch from the Python package index:
pip install datapatch
Example
Given a YAML file like this:
countries:
normalize: true
lowercase: true
asciify: true
options:
- match: Frankreich
value: France
- match:
- Northkorea
- Nordkorea
- Northern Korea
- NKorea
- DPRK
value: North Korea
- contains: Britain
value: Great Britain
The file can be used to apply the data patches against raw input:
from datapatch import read_lookups, LookupException
lookups = read_lookups("countries.yml")
countries = lookups.get("countries")
# This will apply the patch or default to the original string if none exists:
for row in iter_data():
raw = row.get("Country")
row["Country"] = countries.get_value(raw, default=raw)
Extended options
There's a host of options available to configure the application of the data patches:
countries:
# If you mark a lookup as required, a value that matches no options will
# throw a `datapatch.exc:LookupException`.
required: true
# Normalisation will remove many special characters, remove multiple spaces
normalize: false
# By default normalize perform transliteration across alphabets (Путин -> Putin)
# set asciify to false if you want to keep non-ascii alphabets as is
asciify: false
options:
- match: Francois
value: France
# This is a shorthand for defining options that have just one `match` and
# one `value` defined:
map:
Luxemborg: Luxembourg
Lux: Luxembourg
Result objects
You can also have more details associated with a result and access them:
countries:
options:
- match: Frankreich
# These can be arbitrary attributes:
label: France
code: FR
This can be accessed as a result object with attributes:
from datapatch import read_lookups, LookupException
lookups = read_lookups("countries.yml")
countries = lookups.get("countries")
result = countries.match("Frankreich")
print(result.label, result.code)
assert result.capital is None, result.capital
License
datapatch is licensed under the terms of the MIT license, which is included as
LICENSE.
Metadata
Release files for datapatch 1.2.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datapatch-1.2.4-py3-none-any.whl | Python 3 | none | any | Details |
Release files / datapatch-1.2.4-py3-none-any.whl
| Download URL | datapatch-1.2.4-py3-none-any.whl |
|---|---|
| Size | 8.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b6c2dae33a6635d6526b122bcd2229f098ef9f833bd60a93f644a10d82dde699
|
|
BLAKE2b-256 checksum How to use checksums |
d8dd1df187bea2546fa7c3de04d34a366ad5a8095febb61468baddf077c8fd73
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 27, 2025.
Transparency log