A library to make data worse
Project description
Complexifier
Make your pandas even worse!
Problem
When teaching students to work with data, an important lesson is how to clean it.
The problem with this is that there are two types of datasets available on the internet:
- Data that is good, but already cleaned
- Data that is not cleaned, but is terrible and incomprehensible
Complexifier solves this problem by allowing you take the former and turn it into a better version of the latter!
Dependencies
pandastyporandom
Installation
complexifier can be installed using pip
pip install complexifier
Documentation
Usage
Once installed you can use complexifier to add mistakes and outliers to your data
This library has several methods available:
create_spag_error
create_spag_error(word: str) -> str
Introduces a 10% chance of a random spelling error in a given word. This function is useful for simulating typos and spelling mistakes in text data.
introduce_spag_error
introduce_spag_error(df: pd.DataFrame, columns=None) -> pd.DataFrame
Applies the create_spag_error function to each string entry in specified columns of a DataFrame, introducing random spelling errors with a 10% probability.
add_or_subtract_outliers
add_or_subtract_outliers(df: pd.DataFrame, columns=None) -> pd.DataFrame
Randomly adds or subtracts values in specified numeric columns at random indices, simulating outliers between 1% and 10% of the rows.
add_standard_deviations
add_standard_deviations(df: pd.DataFrame, columns=None, min_std=1, max_std=5) -> pd.DataFrame
Adds between 1 to 5 standard deviations to random entries in specified numeric columns to simulate data anomalies.
duplicate_rows
duplicate_rows(df: pd.DataFrame, sample_size=None) -> pd.DataFrame
Introduces duplicate rows into a DataFrame. This function is useful for testing deduplication processes.
add_nulls
add_nulls(df: pd.DataFrame, columns=None, min_percent=1, max_percent=10) -> pd.DataFrame
Inserts null values into specified DataFrame columns. This simulates missing data conditions.
mess_it_up
mess_it_up(df: pd.DataFrame, columns=None, min_std=1, max_std=5, sample_size=None,min_percent=1, max_percent=10, introduce_spag=True, add_outliers=True, add_std=True, duplicate=True, add_null=True) -> pd.DataFrame
Adds all (or some) of the above methods. Really messes it up.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file complexifier-1.0.0.tar.gz.
File metadata
- Download URL: complexifier-1.0.0.tar.gz
- Upload date:
- Size: 443.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
541548994eeade00ea1198679ca163a6298435457ed5a9df602c5d605cb6b673
|
|
| MD5 |
b12a9468800fdf97f06fcdf45a90a77f
|
|
| BLAKE2b-256 |
d5a4831346ed3421dbd5bc9bd0e7a919bfbbf7f22e393202b88363ac70c282e6
|
File details
Details for the file complexifier-1.0.0-py3-none-any.whl.
File metadata
- Download URL: complexifier-1.0.0-py3-none-any.whl
- Upload date:
- Size: 7.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b0239d8323d56a3a4f2d1612ba75b64331906a2322353508fe03ca51091f40ee
|
|
| MD5 |
d609f41eb914df64199f36c1d1b76b2a
|
|
| BLAKE2b-256 |
31738737e69b352a7326211e5fbd40f003fe3a082610cbabc0628b5e0137d330
|