Skip to main content

Flexible framework for transforming Pandas DataFrames using a modular pipeline approach.

Project description

⚙️ DataMorph

Unit Tests Code style: black

Overview

DataMorph is a Python library that provides a flexible framework for transforming Pandas DataFrames using a modular pipeline approach. Transformations are defined in a YAML configuration, and are applied sequentially to your dataset.

By leveraging DataMorph, your pipelines become cleaner, more scalable and easier to debug.

Features

  • Modular and extensible transformation framework.
  • Easily configurable via YAML files.
  • Supports multiple transformations, including:
    • CreateColumn: Creates a new column with a constant value.
    • ColumnsOperator: Performs a math operation on two columns and stores the result in a new column.
    • NormalizeColumn: Applies Z-score normalization.
    • RemoveColumns: Drops specified columns.
    • FillNA: Replaces missing values with a default.
    • MergeDataFrames: Merges two DataFrames based on common keys.
    • And more!
  • Supports custom transformations, defined by the user.

Installation

To install DataMorph in your project:

pip install git+https://github.com/davideganna/DataMorph.git

If you're developing locally:

git clone https://github.com/davideganna/DataMorph.git
cd datamorphers
pip install -e .

Usage

1. Define your initial DataFrame

import pandas as pd

# Sample DataFrame
df = pd.DataFrame(
  {
      'item': ['apple', 'TV', 'banana', 'pasta', 'cake'],
      'item_type': ['food', 'electronics', 'food', 'food', 'food'],
      'price': [3, 100, 2.5, 3, 15],
      'discount_pct': [0.1, 0.05, np.nan, 0.12, np.nan],
  }
)

print(df)
item item_type price discount_pct
apple food 3 0.1
TV electronics 100 0.05
banana food 2.5 nan
pasta food 3 0.12
cake food 15 nan

2. Define Your Transformation Pipeline

Imagine that we want to perform some actions on the original DataFrame. Specifically, we want to identify which items are food, and then calculate the price after a discount percentage is applied. After these operations, we want to polish the DataFrame by removing non interesting columns.

To do so, we create a YAML file specifying a pipeline of transformations, named config.yaml:

pipeline_food:
  - CreateColumn:
      column_name: food_marker
      value: food

  - FilterRows:
      first_column: item_type
      second_column: food_marker
      logic: e

  - FillNA:
      column_name: discount_pct
      value: 0

  - ColumnsOperator:
      first_column: price
      second_column: discount_pct
      logic: mul
      output_column: discount_amount

  - ColumnsOperator:
      first_column: price
      second_column: discount_amount
      logic: sub
      output_column: discounted_price

  - RemoveColumns:
      columns_name:
        - discount_amount
        - food_marker

3. Apply the transformations as defined in the config

Running the pipeline is very simple:

from datamorphers.pipeline_loader import get_pipeline_config, run_pipeline

# Load YAML config
config = get_pipeline_config("config.yaml", pipeline_name='pipeline_food'))

# Run pipeline
transformed_df = run_pipeline(df, config)

print(transformed_df)
item item_type price discount_pct discounted_price
apple food 3 0.1 2.7
banana food 2.5 0 2.5
pasta food 3 0.12 2.64
cake food 15 0 15

Define runtime values in the YAML configuration

DataMorph is flexible, since it can work with variables at runtime:

pipeline_runtime:
  - CreateColumn:
      column_name: ${custom_column_name}
      value: ${custom_value}

Simply pass the arguments you need when you instantiate the pipeline:

custom_column_name = "D"
custom_value = 888

kwargs = {
  "custom_column_name": custom_column_name,
  "custom_value": custom_value
}

config = get_pipeline_config(
    yaml_path=YAML_PATH,
    pipeline_name="pipeline_runtime",
    **kwargs,
)

df = run_pipeline(df, config=config)

Extending datamorphers with Custom Implementations

The datamorphers package allows you to define custom transformations by implementing your own DataMorphers. These user-defined implementations extend the base ones and can be used seamlessly within the pipeline.

Creating a Custom DataMorpher

To define a custom transformation, create a custom_datamorphers.py file in your project and implement a new class that follows the DataMorpher structure:

import pandas as pd
from datamorphers.base import DataMorpher

class CustomTransformer(DataMorpher):
    def __init__(self, column_name: str, value: float):
        self.column_name = column_name
        self.value = value

    def _datamorph(self, df: pd.DataFrame) -> pd.DataFrame:
        """
        Implement your custom transformation here!
        """
        df[self.column_name] = self.value * 3.14
        return df

Importing Custom DataMorphers

To use your custom implementations, create a file named custom_datamorphers.py inside your current directory.

The pipeline will first check for the specified DataMorpher in custom_datamorphers. If it's not found, it will fall back to the default ones in datamorphers. This allows for seamless extension without modifying the base package.

Running the Pipeline with Custom DataMorphers

When defining a pipeline configuration (e.g., in a YAML file), simply reference your custom DataMorpher as you would with a base one:

custom_pipeline:
  CustomTransformer:
    column_name: price
    value: 1.3

Then, execute the pipeline as usual:

df_transformed = run_pipeline(df, config)

If a custom module is provided, your custom transformations will be used instead of (or in addition to) the built-in ones.


Pre-commit Hooks

To ensure code quality, install and configure pre-commit hooks:

pre-commit install
pre-commit run --all-files

Contributing

Contributions are welcome! Please open an issue or submit a pull request.

License

MIT License. See LICENSE for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datamorphers-0.1.tar.gz (11.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datamorphers-0.1-py3-none-any.whl (8.8 kB view details)

Uploaded Python 3

File details

Details for the file datamorphers-0.1.tar.gz.

File metadata

  • Download URL: datamorphers-0.1.tar.gz
  • Upload date:
  • Size: 11.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for datamorphers-0.1.tar.gz
Algorithm Hash digest
SHA256 3d6c2041a0df4684010ef2e3a0f819e3bdeaf8aa531d00c51cd23a26529cd7b6
MD5 379791959149bf79cab7a1a0296ed454
BLAKE2b-256 76a749b7dc4ed018a9e003d2c5d86e7b26d7bcec621bc9cea283994b2d6a2db4

See more details on using hashes here.

File details

Details for the file datamorphers-0.1-py3-none-any.whl.

File metadata

  • Download URL: datamorphers-0.1-py3-none-any.whl
  • Upload date:
  • Size: 8.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.8

File hashes

Hashes for datamorphers-0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 20c0873db2bf1df88698569ae0abb5dff0d596eca088f8059380fff54e011d22
MD5 0de2d9cd60edf9bf5db2b939e5b62fb5
BLAKE2b-256 a465c3947e67b7b84eb21af56ecc72271021f6c9ffea281a1106fddcf786e366

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page