Skip to main content

Automated data profiling and exploratory analysis for tabular datasets.

Project description

📊 Autoprofiler

Autoprofiler is a lightweight, heuristic-based data profiling tool for pandas DataFrames. It generates an HTML report with summaries, correlations, outlier detection, visualizations, and actionable suggestions — designed to scale to large datasets via sampling.

✨ Features

📌 Dataset Overview

Rows and columns

Numeric vs categorical features

Missing data indicators

High-cardinality detection

ML suitability hints

📈 Statistical Summaries

Numeric summaries (count, mean, std, quantiles, missing %)

Categorical summaries (unique values, top category, frequency)

🔗 Correlations

Numeric ↔ Numeric: Pearson & Spearman

Categorical ↔ Categorical: Cramér’s V (with safety heuristics)

Categorical ↔ Numeric: Correlation Ratio (η²)

Automatic skipping of unstable or misleading correlations

Warnings when sampling or heuristics are applied

🚨 Outlier Detection

IQR-based detection

Reports row indices per column

Skips constant or invalid columns

🧠 Heuristic Suggestions

Prioritized recommendations:

🚨 High priority

⚠️ Medium priority

ℹ️ Low priority

Examples:

High skew → consider log/Box-Cox transform

High missingness → consider imputation

Constant columns → consider dropping

📊 Visualizations

Histograms and distributions

Embedded directly into the HTML report (no temp files)

Filename-safe handling for special characters

⚡ Scales to Large Datasets

Automatic sampling for large DataFrames

Sampling warnings included in report

Designed to avoid O(n²) pitfalls where possible

🖥 Example Output

Autoprofiler generates a single self-contained HTML report that includes:

Dataset overview

Tables (summaries & correlations)

Suggestions with priority levels

Embedded plots

Example:

report.html

🚀 Installation

Clone the repository and install dependencies:

``cmd git clone https://github.com/yourusername/autoprofiler.git cd autoprofiler pip install -r requirements.txt


### 🧪 Usage
```python
import pandas as pd
from autoprofiler import profiler

df = pd.read_csv("data.csv")

report = profiler.profile(df)
report.to_html("report.html")

Open report.html in your browser.

📦 Large Dataset Testing

Autoprofiler supports large datasets via sampling.

Example synthetic dataset:

import pandas as pd import numpy as np

n = 1_000_000 df_large = pd.DataFrame({ f"num_{i}": np.random.randn(n) for i in range(10) }) for j in range(5): df_large[f"cat_{j}"] = np.random.choice(["A", "B", "C", "D"], size=n)

df_large.to_csv("synthetic_large.csv", index=False)

df = pd.read_csv("synthetic_large.csv") report = profiler.profile(df) report.to_html("large_report.html")

🧠 Design Philosophy

Transparency over magic

Heuristics over black boxes

Actionable insights over raw statistics

Graceful degradation on large data

This is not meant to replace full EDA — it’s meant to accelerate it.

⚠️ Known Limitations

Correlation measures are heuristic-based

ML readiness scoring is qualitative

Sampling may hide rare patterns

Not intended for real-time profiling

🛠 Roadmap

Execution timing per report section

ML readiness scoring (0–100)

Optional target leakage detection

Stability checks (sample vs full dataset)

Performance optimizations for >10M rows

📄 License

MIT License

🙌 Acknowledgments

Built with:

pandas

numpy

scipy

matplotlib

🧠 Why This Exists

This project was built to:

Understand how data profiling tools work internally

Explore scalability tradeoffs

Practice building interpretable, user-focused data tooling

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

df_autoprofiler-0.1.0.tar.gz (12.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

df_autoprofiler-0.1.0-py3-none-any.whl (13.5 kB view details)

Uploaded Python 3

File details

Details for the file df_autoprofiler-0.1.0.tar.gz.

File metadata

  • Download URL: df_autoprofiler-0.1.0.tar.gz
  • Upload date:
  • Size: 12.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.5

File hashes

Hashes for df_autoprofiler-0.1.0.tar.gz
Algorithm Hash digest
SHA256 df5216e4ed17e9074c7eb9dcc6fb0d926c3c100e281ae97ce57063d0fb1fa7c9
MD5 e5055bf270f18de729b584fec6bb0c23
BLAKE2b-256 22259766512346848bea5b38532965df46c2bbed5aeb545429f001f3f05ae38d

See more details on using hashes here.

File details

Details for the file df_autoprofiler-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for df_autoprofiler-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 07aa1af286eafe9d8e988fe04f99ae024bcb076dbdf2eca70d00b0dccbfc86eb
MD5 57d5e6aa53ad608e2e944019949bd9de
BLAKE2b-256 829ea51cf63a79f45ca1acb60986fc5bd8c6a26ebcc221752eb8b62d758c4474

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page