Skip to main content

Automated data profiling and exploratory analysis for tabular datasets.

Project description

df-autoprofiler

df-autoprofiler is a lightweight, heuristic-based data profiling tool for pandas DataFrames.
It generates a self-contained HTML report with summaries, correlations, outlier detection, visualizations, and actionable suggestions — designed to scale to large datasets via sampling.


1. Dataset Overview

Provides a high-level summary of your dataset:

  • Number of rows and columns
  • Numeric vs categorical feature detection
  • Missing data indicators
  • High-cardinality detection
  • Machine Learning suitability hints

Once you create a profile with from df_autoprofiler.profiler import profile, you can do:

report = profile(df)

You can then access:

  • report.numerics — numeric feature summaries

  • report.categoricals — categorical feature summaries

  • report.correlations — correlation tables

  • report.outliers — outlier detection results

  • report.missing — missing data summaries

  • report.suggestions — actionable suggestions

2. Statistical Summaries

Numeric features: count, mean, std, quantiles, missing %

Categorical features: unique values, top category, frequency

All summaries are available via the report object:

numeric_summary = report.numerics
categorical_summary = report.categoricals

3. Correlations

Numeric ↔ Numeric: Pearson & Spearman

Categorical ↔ Categorical: Cramér’s V (with safety heuristics)

Categorical ↔ Numeric: Correlation Ratio (η²)

Automatic skipping of unstable or misleading correlations

Warnings when sampling or heuristics are applied

Access correlations:

report.correlations["numeric"]["pearson"]
report.correlations["numeric"]["spearman"]
report.correlations["mixed"]

4. Outlier Detection

IQR-based detection

Reports row indices per column

Skips constant or invalid columns

Access outliers:

outliers = report.outliers

5. Missing Data Analysis

Detects missing values across columns

Generates summary tables with counts and percentages

Access missing data:

missing_data = report.missing

6. Heuristic Suggestions

Prioritized recommendations with three levels:

  • 🚨 High priority

  • ⚠️ Medium priority

  • ℹ️ Low priority

Examples:

High skew → consider log/Box-Cox transform

High missingness → consider imputation

Constant columns → consider dropping

Access suggestions:

suggestions = report.suggestions

7. Visualizations

Histograms, distributions, and correlation plots

Embedded directly into the HTML report

Handles special characters safely in filenames

Generate plots:

report.plots['hist_{NAME_OF_COL}']

Quickstart Example

import pandas as pd
from df_autoprofiler.profiler import profile

df = pd.read_csv("data.csv")

# Create the profile
report = profile(df)

# Access summaries
print(report.numerics.head())
print(report.categoricals.head())
print(report.missing.head())

# Generate HTML report
report.to_html("report.html")

Handling Large Datasets

  • Automatic sampling to maintain performance

  • Warnings included in report when sampling is applied

Example with a large synthetic dataset:

import pandas as pd
import numpy as np
from df_autoprofiler.profiler import profile

n = 1_000_000
df_large = pd.DataFrame({
    f"num_{i}": np.random.randn(n) for i in range(10)
})
for j in range(5):
    df_large[f"cat_{j}"] = np.random.choice(["A","B","C","D"], size=n)

report = profile(df_large)
report.to_html("large_report.html")

Installation

pip install df-autoprofiler

Known Limitations

  • Correlation measures are heuristic-based

  • ML readiness scoring is qualitative

  • Sampling may hide rare patterns

  • Bugs

📄 License

MIT License

🙌 Acknowledgments

Built with:

  • pandas

  • numpy

  • scipy

  • matplotlib

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

df_autoprofiler-0.1.1.tar.gz (12.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

df_autoprofiler-0.1.1-py3-none-any.whl (13.4 kB view details)

Uploaded Python 3

File details

Details for the file df_autoprofiler-0.1.1.tar.gz.

File metadata

  • Download URL: df_autoprofiler-0.1.1.tar.gz
  • Upload date:
  • Size: 12.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.5

File hashes

Hashes for df_autoprofiler-0.1.1.tar.gz
Algorithm Hash digest
SHA256 f5ea14ab442bc534361c3f9772707f9de01277b76b2fb13ca2080821b01a265f
MD5 e1f24a6100467b554a0f59f64511d20c
BLAKE2b-256 7c3efcc4222145ff6e1d3894dc55535ec068430daba340b33ee30af400df43a1

See more details on using hashes here.

File details

Details for the file df_autoprofiler-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for df_autoprofiler-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 6fc61291a019704f0e91c0a9ab362c6c382b970bebac1f6ea730e945acdf5112
MD5 18a3d16aa1de691f6a4d37c9ffdef8f8
BLAKE2b-256 114b02cd7fc186ecab368db1be967619f275dcc6297224ff9f29ef2b910d1b8e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page