A DataFrame workflow accelerator for fast, deterministic data prep and analysis.

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

sayantanpatra

These details have not been verified by PyPI

Project description

dfxpy: The DataFrame Workflow Accelerator ⚡

dfxpy Logo

dfxpy is a high-performance Python library designed to reduce the entire data preparation and analysis workflow into a few powerful, deterministic operations. It helps data scientists and engineers go from raw, messy datasets to model-ready features in seconds, not hours.

🎯 Why dfxpy?

While pandas is the industry standard for data manipulation, cleaning a dataset for machine learning often requires writing repetitive boilerplate code. dfxpy acts as an accelerator, automating the "grunt work" while providing intelligent diagnostics and insights.

Fast & Deterministic: Optimized operations without the unpredictability of AI.
Production-Ready: Type-hinted, modular, and follows PEP8 standards.
Lightweight: Minimal dependencies—built on top of pandas, numpy, and scikit-learn.
Self-Contained: Load and analyze data without needing to import multiple libraries.

✨ Key Features

🛠️ Auto-Pilot Cleaning (`auto`)

One-line pipeline for column normalization (snake_case), duplicate removal, smart type inference (Object → Numeric/Datetime), and missing value imputation.

🔍 Deep Audit (`audit`)

Intelligent dataset diagnostics that detect:

ID-like columns (uniqueness checks)
High cardinality categoricals
Multicollinearity (correlated features)
Data skewness and low-variance columns

🚀 ML Preparation (`prepare`)

Transform your data into X and y instantly. Automatically handles categorical target encoding (LabelEncoding), feature encoding (One-Hot), and optional scaling.

📉 Smart EDA (`eda`)

Generates structured, human-readable summaries of shapes, nulls, unique counts, and correlation matrices.

🧹 Outlier Management (`outliers`)

Detect and handle outliers using the IQR (Interquartile Range) method with options to remove or cap.

📊 One-Click Reports (`report`)

Generate beautiful, standalone HTML EDA reports with distributions and missing value summaries—zero dependencies required.

🔄 Dataset Comparison (`compare`)

Instantly track changes between two versions of a dataset. Detects shape shifts, column changes, and cell-level value modifications.

⚖️ Smart Balancing (`balance`)

Handle imbalanced datasets with ease using Oversampling, Undersampling, or synthetic interpolation.

🛡️ Data Contracts (`validate`)

Enforce rigorous data contracts with schema validation. Ensure your pipeline only processes data that meets your quality standards.

🏗️ Workflow Pipelines (`pipeline`)

Build, save, and reuse complex cleaning workflows with the new chainable pipeline engine.

💡 ML Model Advisor (`suggest`)

Get instant recommendations on which ML algorithms (XGBoost, LightGBM, etc.) to use based on your dataset's specific characteristics.

🕵️ Explainable Cleaning (`analyze_cleaning`)

Understand exactly what happened to your data. Logs every transformation and repair for full transparency.

📦 Installation

Install the latest version via pip:

pip install dfxpy

For detailed usage and parameter definitions, see the Full API Reference.

🚀 Quick Start

Python API

import dfxpy as dfx

# 1. Load data directly
df = dfx.load("raw_data.csv")

# 2. Run auto-clean (names, types, nulls, encoding)
df_clean = dfx.auto(df)

# 3. Get intelligent insights
dfx.audit(df_clean)

# 4. Prepare for Machine Learning
X, y = dfx.prepare(df_clean, target="outcome")

Command Line Interface (CLI)

# Analyze a dataset instantly
dfxpy analyze data.csv

# Prepare data for ML via CLI
dfxpy prepare data.csv --target price --output cleaned_features.csv

🧠 Pandas vs. dfxpy

Task	Standard Pandas	dfxpy
Load Data	`pd.read_csv("data.csv")`	`dfx.load("data.csv")`
Clean Names	`df.columns = [c.lower().replace(' ', '_') for c in df.columns]`	`dfx.auto(df)`
Handle Nulls	`df['val'].fillna(df['val'].median())`	`dfx.auto(df)`
ML Prep	10+ lines (Split, Encode, Scale)	`dfx.prepare(df, target='y')`
Audit	Manual inspection	`dfx.audit(df)`
HTML Report	50+ lines (Matplotlib/Jinja)	`dfx.report(df)`
Dataset Diff	`df1.compare(df2)` (limited)	`dfx.compare(df1, df2)`
Explainability	Manual logging	`dfx.analyze_cleaning(df)`

🤝 Contributing

We welcome contributions! Please see our CONTRIBUTING.md for guidelines on how to submit issues, feature requests, and pull requests.

📄 License

dfxpy is licensed under the MIT License.

Built with ❤️ for the Data Science Community.

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

sayantanpatra

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.3.5

May 6, 2026

This version

0.3.4

May 6, 2026

0.3.3

May 6, 2026

0.3.2

May 6, 2026

0.3.1

May 6, 2026

0.3.0

May 6, 2026

0.2.6

May 5, 2026

0.2.5

May 5, 2026

0.2.4

May 5, 2026

0.2.3

May 5, 2026

0.2.2

May 5, 2026

0.2.1

May 5, 2026

0.2.0

May 5, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dfxpy-0.3.4.tar.gz (25.5 kB view details)

Uploaded May 6, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

dfxpy-0.3.4-py3-none-any.whl (27.3 kB view details)

Uploaded May 6, 2026 Python 3

File details

Details for the file dfxpy-0.3.4.tar.gz.

File metadata

Download URL: dfxpy-0.3.4.tar.gz
Upload date: May 6, 2026
Size: 25.5 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for dfxpy-0.3.4.tar.gz
Algorithm	Hash digest
SHA256	`1927941e95283282e09a3688e35f7b0da40c58057376fad0bc25d052cd871160`
MD5	`b78e31ec57e55a3d9dea7ba58ba9fa6a`
BLAKE2b-256	`0b771466419a5a1c41cf95f308285eacc92bcce4ef00efda924a2cf076801e1e`

See more details on using hashes here.

Provenance

The following attestation bundles were made for dfxpy-0.3.4.tar.gz:

Publisher: publish.yml on sayantancodex/dfxpy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: dfxpy-0.3.4.tar.gz
- Subject digest: 1927941e95283282e09a3688e35f7b0da40c58057376fad0bc25d052cd871160
- Sigstore transparency entry: 1449714171
- Sigstore integration time: May 6, 2026
Source repository:
- Permalink: sayantancodex/dfxpy@e6da9e4741411926b4e8bb0b7a31a451f4aabd8a
- Branch / Tag: refs/tags/v0.3.4
- Owner: https://github.com/sayantancodex
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@e6da9e4741411926b4e8bb0b7a31a451f4aabd8a
- Trigger Event: push

File details

Details for the file dfxpy-0.3.4-py3-none-any.whl.

File metadata

Download URL: dfxpy-0.3.4-py3-none-any.whl
Upload date: May 6, 2026
Size: 27.3 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for dfxpy-0.3.4-py3-none-any.whl
Algorithm	Hash digest
SHA256	`294451dac69c03c67d07a3abffea64bb644b63f057bc7747d3e519577d06eb75`
MD5	`0a92338cb469029e2fd04a9daa82316c`
BLAKE2b-256	`2fb889043ae540fa3417a0f5e734a9be1a10104baf5d19dc442d5481f023890f`

See more details on using hashes here.

Provenance

The following attestation bundles were made for dfxpy-0.3.4-py3-none-any.whl:

Publisher: publish.yml on sayantancodex/dfxpy

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: dfxpy-0.3.4-py3-none-any.whl
- Subject digest: 294451dac69c03c67d07a3abffea64bb644b63f057bc7747d3e519577d06eb75
- Sigstore transparency entry: 1449714189
- Sigstore integration time: May 6, 2026
Source repository:
- Permalink: sayantancodex/dfxpy@e6da9e4741411926b4e8bb0b7a31a451f4aabd8a
- Branch / Tag: refs/tags/v0.3.4
- Owner: https://github.com/sayantancodex
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@e6da9e4741411926b4e8bb0b7a31a451f4aabd8a
- Trigger Event: push

dfxpy 0.3.4

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Project description

dfxpy: The DataFrame Workflow Accelerator ⚡

🎯 Why dfxpy?

✨ Key Features

🛠️ Auto-Pilot Cleaning (auto)

🔍 Deep Audit (audit)

🚀 ML Preparation (prepare)

📉 Smart EDA (eda)

🧹 Outlier Management (outliers)

📊 One-Click Reports (report)

🔄 Dataset Comparison (compare)

⚖️ Smart Balancing (balance)

🛡️ Data Contracts (validate)

🏗️ Workflow Pipelines (pipeline)

💡 ML Model Advisor (suggest)

🕵️ Explainable Cleaning (analyze_cleaning)

📦 Installation

🚀 Quick Start

Python API

Command Line Interface (CLI)

🧠 Pandas vs. dfxpy

🤝 Contributing

📄 License

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance

🛠️ Auto-Pilot Cleaning (`auto`)

🔍 Deep Audit (`audit`)

🚀 ML Preparation (`prepare`)

📉 Smart EDA (`eda`)

🧹 Outlier Management (`outliers`)

📊 One-Click Reports (`report`)

🔄 Dataset Comparison (`compare`)

⚖️ Smart Balancing (`balance`)

🛡️ Data Contracts (`validate`)

🏗️ Workflow Pipelines (`pipeline`)

💡 ML Model Advisor (`suggest`)

🕵️ Explainable Cleaning (`analyze_cleaning`)