dataset-doctor
Automatic Dataset Diagnosis and Cleaning for Machine Learning
dataset-doctor helps you quickly identify and fix common dataset quality problems before model training.
It is built for:
- Data scientists preparing tabular data for experiments
- ML engineers standardizing preprocessing workflows
- Beginners who want safer defaults for dataset cleaning
Instead of manually writing repeated preprocessing code, you can diagnose data issues and run an automatic cleaning pipeline from either Python or the command line.
Features
- Dataset diagnosis with a readable summary report
- Missing value detection and imputation
- Duplicate row detection and removal
- Outlier detection and handling
- Constant column detection and removal
- Optional normalization for numeric columns
- YAML-based configuration system for preprocessing behavior
- CLI commands for diagnosis, cleaning, display, and config generation
Installation
Install from PyPI:
pip install dataset-doctor
Install from source:
git clone https://github.com/Mirdula18/dataset-doctor.git
cd dataset-doctor
pip install .
Quick Example
import dataset_doctor as dd
report = dd.diagnose("data.csv")
print(report.summary())
clean_df = dd.auto_fix("data.csv")
CLI Usage
Diagnose a dataset:
dataset-doctor diagnose data.csv
Print report output (alias of diagnose):
dataset-doctor report data.csv
Clean a dataset:
dataset-doctor clean data.csv
Clean and write output file:
dataset-doctor clean data.csv --output cleaned.csv
Enable normalization from CLI:
dataset-doctor clean data.csv --normalize
Clean using a YAML config file:
dataset-doctor clean data.csv --config dataset_doctor_config.yaml
Generate a default config file:
dataset-doctor init-config
Display rows:
dataset-doctor display data.csv --rows 10
Show rows (alias of display):
dataset-doctor show data.csv --tail --rows 20 --columns age,salary
Python API
Main API entry points:
- dd.diagnose(dataset)
- dd.auto_fix(dataset, ...)
- dd.display_data(dataset, ...)
Example Usecase
import dataset_doctor as dd
dd.diagnose("data.csv")
dd.auto_fix("data.csv")
dd.auto_fix("data.csv", output_path="cleaned.csv")
dd.auto_fix("data.csv", output="cleaned.csv")
dd.auto_fix("data.csv", do_normalize=True)
dd.auto_fix("data.csv", return_scaler=True)
dd.auto_fix("data.csv", config="dataset_doctor_config.yaml")
dd.auto_fix("data.csv", config={"missing_values": {"numeric_strategy": "mean"}})
dd.display_data("data.csv")
dd.display_data("data.csv", rows=10)
dd.display_data("data.csv", tail=True)
dd.display_data("data.csv", columns=["col1", "col2"])
dd.display_data("data.csv", all_rows=True)
report = dd.diagnose("data.csv")
report.summary()
report.to_dict()
report.print_report()
Example:
import dataset_doctor as dd
# Diagnose
report = dd.diagnose("data.csv")
print(report.summary())
# Auto-clean with options
clean_df = dd.auto_fix(
"data.csv",
output_path="cleaned.csv",
do_normalize=True,
)
Configuration System
Use a YAML file to customize preprocessing behavior.
CLI example:
dataset-doctor clean data.csv --config config.yaml
Python example:
import dataset_doctor as dd
clean_df = dd.auto_fix("data.csv", config="config.yaml")
Example config:
missing_values:
numeric_strategy: median
categorical_strategy: mode
max_missing_threshold: 0.4
duplicates:
remove: true
outliers:
method: iqr
action: clip
normalization:
method: minmax
range: [0, 1]
feature_selection:
remove_constant_columns: true
correlation_threshold: 0.9
logging:
verbosity: medium
Output Example
## DATASET DIAGNOSIS REPORT
Rows: 10000
Columns: 12
### Issues Detected
Missing Values:
- age (12.0%)
- salary (4.0%)
Duplicate Rows:
- 18 rows
Outliers:
- transaction_amount (42 values)
Constant Columns:
- user_flag
Highly Correlated Columns:
- income vs salary (0.97)
License
MIT License. See LICENSE for details.
Release files for dataset-doctor 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dataset_doctor-1.0.1.tar.gz | 22.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dataset_doctor-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 46.7 kB
Release files / dataset_doctor-1.0.1.tar.gz
| Download URL | dataset_doctor-1.0.1.tar.gz |
|---|---|
| Size | 22.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
97cb93b06d9e19976005063b2720e6ae6a0c571871912d4228f1c9cea6118a2d
|
|
BLAKE2b-256 checksum How to use checksums |
58f35f911e6405c7d9d6ac865917ae2ba48efcef0fd4954e9174f5fa6f755744
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.15
|
Release files / dataset_doctor-1.0.1-py3-none-any.whl
| Download URL | dataset_doctor-1.0.1-py3-none-any.whl |
|---|---|
| Size | 24.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8f65465c5032fac926ba8c0a006502f91052a55614a475fe9a89fed206013e2b
|
|
BLAKE2b-256 checksum How to use checksums |
03a18734eeb7a198f3a70e2bd9f6bbcb0958a13f55bd929cc0840b19d6050332
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.15
|