Automatic feature engineering, insights, and sklearn pipelines for tabular ML with optional time-series support.
Project description
FeatureCraft
Automatic feature engineering with beautiful explanations - zero configuration, maximum transparency!
Intelligent preprocessing pipelines + automatic explanations for every decision + sklearn integration
Quick Start
pip install featurecraft
from featurecraft.pipeline import AutoFeatureEngineer
import pandas as pd
# Load data
df = pd.read_csv("data.csv")
X, y = df.drop("target", axis=1), df["target"]
# Fit and transform with automatic explanations!
afe = AutoFeatureEngineer()
X_transformed = afe.fit_transform(X, y, estimator_family="tree")
# Beautiful formatted explanations automatically print showing:
# - Why each transformation was chosen
# - What columns were affected
# - Configuration parameters used
# - Performance tips and recommendations
That's it! Zero configuration, maximum transparency.
Table of Contents
- Quick Start
- About The Project
- Getting Started
- Usage
- Examples
- Documentation
- Roadmap
- Contributing
- License
About The Project
FeatureCraft is a comprehensive feature engineering library that automates the process of transforming raw tabular data into machine learning-ready features. It provides intelligent preprocessing, feature selection, encoding, and scaling tailored for different estimator families.
Key Features
✨ Automatic Feature Engineering: Intelligent preprocessing pipeline that handles missing values, outliers, categorical encoding, and feature scaling
🔮 Automatic Explainability: Beautiful, automatic explanations of every transformation decision printed to console - understand WHY each preprocessing choice was made without any extra code!
🔍 Dataset Analysis: Comprehensive insights into your data including distributions, correlations, and data quality issues
📊 Multiple Estimator Support: Optimized preprocessing for tree-based models, linear models, SVMs, k-NN, and neural networks
🛠️ Sklearn Integration: Seamless integration with scikit-learn pipelines and ecosystem
📈 HTML Reports: Interactive visualizations and insights reports
⚡ CLI & Python API: Choose between command-line interface or programmatic usage
🕐 Time Series Support: Optional time-series aware preprocessing
🤖 AI-Powered Intelligence: Advanced LLM-driven feature engineering with adaptive optimization
AI-Powered Intelligence
FeatureCraft includes sophisticated AI components that leverage Large Language Models and machine learning to make intelligent feature engineering decisions:
🤖 AIFeatureAdvisor
Uses Large Language Models (OpenAI, Anthropic, or local models) to analyze your dataset characteristics and recommend optimal feature engineering strategies. It considers:
- Column types and distributions
- Missing value patterns
- Outlier characteristics
- Feature interactions and relationships
- Domain-specific preprocessing requirements
📋 FeatureEngineeringPlanner
Orchestrates the entire feature engineering workflow:
- Analyzes dataset characteristics
- Gets AI recommendations for optimal strategies
- Applies smart optimizations based on data patterns
- Configures preprocessing pipelines automatically
🔄 AdaptiveConfigOptimizer
Learns from model performance feedback to continuously improve feature engineering strategies:
- Tracks performance metrics across different datasets
- Identifies successful patterns and configurations
- Adapts recommendations based on historical results
- Prevents overfitting through intelligent feature selection
Why FeatureCraft?
- Automated Workflow: No need to manually handle different data types and preprocessing steps
- Automatic Explanations: See beautiful formatted explanations of every decision - enabled by default with zero extra code!
- Best Practices: Implements proven feature engineering techniques
- Performance Optimized: Different preprocessing strategies for different model types
- Production Ready: Exports sklearn-compatible pipelines for deployment
- Comprehensive Analysis: Deep insights into your dataset characteristics
- Explainable AI: Understand why transformations are applied and how they affect your data
- Transparency: Rich explanations help build trust and debug preprocessing decisions
Getting Started
Prerequisites
- Python 3.9 or higher
- pandas >= 1.5
- scikit-learn >= 1.3
- numpy >= 1.23
Installation
pip install featurecraft
All features including AI-powered planning, enhanced encoders, SHAP explainability, and schema validation are included by default.
Usage
CLI Usage
# Analyze dataset and generate comprehensive report
featurecraft analyze --input data.csv --target target_column --out artifacts/
# Fit preprocessing pipeline and transform data
featurecraft fit-transform --input data.csv --target target_column --out artifacts/ --estimator-family tree
# Open the generated HTML report
open artifacts/report.html
Python API
Basic Usage with Automatic Explanations
import pandas as pd
from featurecraft.pipeline import AutoFeatureEngineer
from featurecraft.config import FeatureCraftConfig
# Load your data
df = pd.read_csv("your_data.csv")
X, y = df.drop(columns=["target_column"]), df["target_column"]
# Initialize with automatic explanations (enabled by default)
config = FeatureCraftConfig(
explain_transformations=True, # Enable explanations (default)
explain_auto_print=True # Auto-print after fit (default)
)
afe = AutoFeatureEngineer(config=config)
# Fit and transform - explanations print automatically!
Xt = afe.fit_transform(X, y, estimator_family="tree")
# Beautiful formatted explanations appear here showing:
# - Column classifications
# - Imputation strategies
# - Encoding decisions
# - Scaling choices
# - Feature transformations
# - And WHY each decision was made!
print(f"Transformed {X.shape[1]} features into {Xt.shape[1]} features")
# Export pipeline for production use
afe.export("artifacts")
Advanced: Analyze + Manual Explanation Control
# Analyze dataset first (optional but recommended)
summary = afe.analyze(df, target="target_column")
print(f"Detected task: {summary.task}")
print(f"Found {len(summary.issues)} data quality issues")
# Fit with manual explanation control
config = FeatureCraftConfig(explain_auto_print=False) # Disable auto-print
afe = AutoFeatureEngineer(config=config)
afe.fit(X, y, estimator_family="linear")
# Print or save explanations manually
afe.print_explanation() # Rich console output
afe.save_explanation("artifacts/explanation.md", format="markdown")
afe.save_explanation("artifacts/explanation.json", format="json")
Estimator Families
Choose the preprocessing strategy based on your model type:
| Family | Models | Scaling | Encoding | Best For |
|---|---|---|---|---|
tree |
XGBoost, LightGBM, Random Forest | None | Label Encoding | Tree-based models |
linear |
Linear/Logistic Regression | StandardScaler | One-hot + Target | Linear models |
svm |
SVM, SVC | StandardScaler | One-hot | Support Vector Machines |
knn |
k-Nearest Neighbors | MinMaxScaler | Label Encoding | Distance-based models |
nn |
Neural Networks | MinMaxScaler | Label Encoding | Deep learning |
Configuration
Customize the preprocessing behavior:
from featurecraft.config import FeatureCraftConfig
# Custom configuration
config = FeatureCraftConfig(
low_cardinality_max=15, # Max unique values for low-cardinality features
outlier_share_threshold=0.1, # Threshold for outlier detection
random_state=42, # For reproducible results
# Explainability (enabled by default!)
explain_transformations=True, # Enable detailed explanations (default: True)
explain_auto_print=True, # Auto-print after fit() (default: True)
explain_save_path=None, # Optional: auto-save to file path
)
afe = AutoFeatureEngineer(config=config)
# Note: You can also use AutoFeatureEngineer() without config
# and get automatic explanations by default!
Output Artifacts
The export() method creates:
pipeline.joblib: Fitted sklearn Pipeline ready for productionmetadata.json: Configuration and processing summaryfeature_names.txt: List of all output feature namesexplanation.md: Human-readable explanation of transformation decisions (when explanations enabled)explanation.json: Machine-readable explanation data (when explanations enabled)
The analyze() method generates:
report.html: Interactive HTML report with plots, insights, and recommendations
Examples
Check out the examples directory for comprehensive usage examples:
- 01_quickstart.py: Basic usage with automatic explanations on Iris and Wine datasets
- 02_kaggle_benchmark.py: Kaggle dataset benchmarking
- 03_complex_kaggle_benchmark.py: Advanced benchmarking with complex datasets
- 06_explainability_demo.py: Deep dive into explanation features and export formats
Run the quickstart example to see automatic explanations in action:
python examples/01_quickstart.py
This will demonstrate automatic feature engineering with beautiful formatted explanations for the Iris and Wine datasets!
Documentation
Getting Started
- Getting Started: Quick start guide for new users
- API Reference: Complete Python API documentation
- CLI Reference: Command-line interface guide
Configuration & Optimization
- Configuration Guide: Comprehensive parameter reference
- Optimization Guide: Performance tuning and best practices
Advanced Topics
- Advanced Features: Drift detection, leakage prevention, SHAP
- Benchmarks: Real-world dataset performance results
- Troubleshooting: Common issues and solutions
Roadmap
See the open issues for a list of proposed features and known issues.
- Enhanced time series preprocessing
- Feature selection algorithms
- Integration with popular ML frameworks
- GPU acceleration support
- Advanced outlier detection methods
- Automated feature interaction detection
Contributing
Contributions are what make the open source community such an amazing place to learn, inspire, and create. Any contributions you make are greatly appreciated.
If you have a suggestion that would make this better, please fork the repo and create a pull request. You can also simply open an issue with the tag "enhancement".
Don't forget to give the project a star! Thanks again!
License
Distributed under the MIT License. See LICENSE for more information.
Built with ❤️ for the machine learning community
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file featurecraft-1.0.0.tar.gz.
File metadata
- Download URL: featurecraft-1.0.0.tar.gz
- Upload date:
- Size: 162.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
afbf026bd66d8318bb6be9d833f9833203dc2feafdcddef6a880fcd612f6858d
|
|
| MD5 |
273ae1637dc1fc032a8cc4d6f10b3f7d
|
|
| BLAKE2b-256 |
bfd41c72e5c1630e825fe348feca16e782e00751687737bb2e4e1661fcb2853b
|
File details
Details for the file featurecraft-1.0.0-py3-none-any.whl.
File metadata
- Download URL: featurecraft-1.0.0-py3-none-any.whl
- Upload date:
- Size: 178.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1682eb33cc3fd003e0686e79f55849ad4d53ca2d53ec649e6b66ab9187e4c42d
|
|
| MD5 |
1c6de8b19dbe584de9dcdff042358e39
|
|
| BLAKE2b-256 |
e90abd81dc4441a5a09ba0de808a7ef59f53700b960b7388a082abb50a907a3d
|