A library to make data worse
Project description
Complexifier
This makes your pandas dataframe even worse
Dependencies
pandastyporandom
Installation
complexifier can be installed using pip
pip install complexifier
Usage
Once installed you can use complexifier to add mistakes and outliers to your data
This library has several methods available:
create_spag_error(word: str) -> str
Introduces a 10% chance of a random spelling error in a given word. This function is useful for simulating typos and spelling mistakes in text data.
introduce_spag_error(df: pd.DataFrame, columns=None) -> pd.DataFrame
Applies the create_spag_error function to each string entry in specified columns of a DataFrame, introducing random spelling errors with a 10% probability.
Parameters:
df: The DataFrame to be altered.columns: Optional; specify column names to apply errors to. If not provided, it defaults to all string columns.
add_or_subtract_outliers(df: pd.DataFrame, columns=None) -> pd.DataFrame
Randomly adds or subtracts values in specified numeric columns at random indices, simulating outliers between 1% and 10% of the rows.
Parameters:
df: DataFrame to be modified.columns: Optional; specify columns to adjust.
add_standard_deviations(df: pd.DataFrame, columns=None, min_std=1, max_std=5) -> pd.DataFrame
Adds between 1 to 5 standard deviations to random entries in specified numeric columns to simulate data anomalies.
Parameters:
df: The DataFrame to manipulate.columns: Optional; specify columns to modify.min_std: Minimum number of standard deviations to add.max_std: Maximum number of standard deviations to add.
duplicate_rows(df: pd.DataFrame, sample_size=None) -> pd.DataFrame
Introduces duplicate rows into a DataFrame. This function is useful for testing deduplication processes.
Parameters:
df: DataFrame where duplicates will be introduced.sample_size: Optional; number of rows to duplicate. A random percentage between 1% and 10% if not specified.
add_nulls(df: pd.DataFrame, columns=None, min_percent=1, max_percent=10) -> pd.DataFrame
Inserts null values into specified DataFrame columns. This simulates missing data conditions.
Parameters:
df: The DataFrame to modify.columns: Optional; specific columns to addnullsto.min_percent: Minimum percentage ofnullvalues to insert.max_percent: Maximum percentage ofnullvalues to insert.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file complexifier-0.2.3.tar.gz.
File metadata
- Download URL: complexifier-0.2.3.tar.gz
- Upload date:
- Size: 4.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.1.1 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8836819749205fd06ca5b390de1e3b7de01473058546832df8adacfd085f69dd
|
|
| MD5 |
d642ac30c313498459d77b3d1d5aad05
|
|
| BLAKE2b-256 |
7c15fef8e8997dfecab7c3cf46df39328170bf66ffd3545fa57f0c9bede35db3
|
File details
Details for the file complexifier-0.2.3-py3-none-any.whl.
File metadata
- Download URL: complexifier-0.2.3-py3-none-any.whl
- Upload date:
- Size: 5.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.1.1 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5817b7cac301be235ab7abe9988680e4a62f7b8b7ae4f9f42a191c3965fd1a51
|
|
| MD5 |
1c60043ee2487abcf168bc49e1b5071f
|
|
| BLAKE2b-256 |
6aaa773c831ee2d10f26a2065deff7329615efd3954a9d746dad45c52012ad14
|