Skip to main content

Data Transformation Library

Project description

Python

Data Transformation Library

This project is a POC for a Python library designed to assist data scientists with data transformations.

Key Features:

This library consists of 3 functions: Transpose, Time Series Windowing and Cross-Correlation. Here's a brief explanation of these functions:

  1. Transpose: Transposes a given matrix (2D tensor).

    • Function Signature:
    def transpose2d(input_matrix: list[list[float]]) -> list[list[float]]:
    
    • Parameters: input_matrix: A list of lists of real numbers, where each inner list represents a row in the matrix.

    • Returns: A new list of lists representing the transposed matrix, where the rows of the original matrix become the columns of the new matrix, and the columns of the original matrix become the rows of the new matrix.

    • Example:

    # Input
    input_matrix = [
    [1.0, 2.0, 3.0],
    [4.0, 5.0, 6.0]
    ]
    
    # Output
    output_matrix = [
    [1.0, 4.0],
    [2.0, 5.0],
    [3.0, 6.0]
    ]
    

    In this example, the input matrix with dimensions 2x3 (2 rows and 3 columns) is transformed into an output matrix with dimensions 3x2 (3 rows and 2 columns).

  2. Window Extraction: Extracts overlapping windows from a 1D array with specified size, shift, and stride.

    • Function Signature:
    def window1d(input_array: list | np.ndarray, size: int, shift: int = 1, stride: int = 1) -> list[list | np.ndarray]:
    
    • Parameters: input_array: A list or 1D Numpy array of real numbers. size: A positive integer that determines the size (length) of each window. shift: A positive integer that determines the shift (step size) between different windows. Default is 1. stride: A positive integer that determines the stride (step size) within each window. Default is 1.

    • Returns: A list of lists or 1D Numpy arrays of real numbers, where each element represents a window extracted from the input array according to the specified parameters.

    • Example:

    # Input
    input_array = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0]
    size = 3
    shift = 2
    stride = 1
    
    # Output
    output_matrix = [
    [1.0, 2.0, 3.0],
    [3.0, 4.0, 5.0]
    ]
    

    In this example, the input array [1.0, 2.0, 3.0, 4.0, 5.0, 6.0] is processed with a window size of 3, a shift of 2, and a stride of 1. This results in extracting windows of length 3 from the input array, where each new window starts 2 elements after the previous one, and each window advances one element at a time.

  3. Cross-Correlation: Applies a 2D convolution operation on a matrix using a given kernel and stride.

    • Function Signature:
    def convolution2d(input_matrix: np.ndarray, kernel: np.ndarray, stride: int = 1) -> np.ndarray:
    
    • Parameters: input_matrix: A 2D Numpy array of real numbers representing the input matrix to be convolved. kernel: A 2D Numpy array of real numbers representing the convolution kernel. stride: An integer greater than 0 that determines the stride (step size) for moving the kernel across the input matrix. Default is 1.

    • Returns: A 2D Numpy array of real numbers representing the result of applying the convolution operation. The output matrix will have reduced dimensions compared to the input matrix, depending on the size of the kernel and the stride.

    • Example:

    # Input
    input_matrix = np.array([
    [1.0, 2.0, 3.0],
    [4.0, 5.0, 6.0],
    [7.0, 8.0, 9.0]
    ])
    kernel = np.array([
    [1.0, 0.0],
    [0.0, -1.0]
    ])
    stride = 1
    
    # Output
    output_matrix = np.array([
    [ 6.0,  6.0],
    [ 0.0, -6.0]
    ])
    

    In this example, the input_matrix is convolved with the kernel using a stride of 1. The resulting output_matrix is a 2D array where each element represents the result of applying the kernel to the corresponding region of the input matrix.

Installation

Follow these steps to initialize and run this Poetry-based project in a new environment.

  1. The library is available on PyPI and can be installed using pip:

    pip install data-transformation-library-jmarci
    

Alternatively: Clone this repository from your version control system (e.g., GitHub, GitLab):

```bash
git clone <repository_url>
cd <repository_name>
```
  1. Ensure that Poetry is installed in your new environment. If not, you can install it using the official installation script:

    curl -sSL https://install.python-poetry.org | python3 -
    
  2. Navigate to the project directory and install the dependencies specified in the pyproject.toml file using Poetry:

    cd <project_directory>
    poetry install
    

    This command will create a virtual environment (if it doesn't already exist) and install all the dependencies listed in pyproject.toml and poetry.lock.

  3. Poetry automatically manages virtual environments for you. To activate the virtual environment, you can use:

    poetry shell
    

    This will activate the environment, and you can start working within it.

  4. After installation, the functions can be imported and used in Python scripts or Jupyter notebooks.

    • Example usage:
    from data_transformation_library.transpose import transpose2d
    from data_transformation_library.time_series_windowing import window1d
    from data_transformation_library.cross_correlation import convolution2d
    import numpy as np
    
    # Example for transpose2d
    matrix = [[1, 2], [3, 4]]
    transposed_matrix = transpose2d(matrix)
    print(transposed_matrix)
    
    # Example for window1d
    input_array = [1, 2, 3, 4, 5]
    size = 3
    shift = 3
    list_of_lists = window1d(input_array, size, shift)
    print(list_of_lists)
    
    # Example for convolution2d
    input_matrix = np.array([[1, 2, 3], [4, 5, 6]])
    kernel = np.array([[1, 0], [0, -1]])
    stride = 1
    np_array = convolution2d(input_matrix, kernel, stride)
    print(np_array)
    

Suggestions for Future Improvements

  • Error Handling: Include robust error handling to manage invalid input types and values. Raise appropriate exceptions for clear feedback.
  • Parameter Validation: Validate parameters to ensure they meet expected criteria (e.g., positive integers, correct dimensions).
  • Testing: Add comprehensive test cases to cover typical and edge cases. Ensure that the functions behave correctly across a range of inputs.

Feel free to fork this repository and make your own modifications!

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

data_transformation_library_jmarci-0.1.3.tar.gz (4.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file data_transformation_library_jmarci-0.1.3.tar.gz.

File metadata

File hashes

Hashes for data_transformation_library_jmarci-0.1.3.tar.gz
Algorithm Hash digest
SHA256 e35f8c7ef964517e1362087df01a36cc66187352562f6929a736971ed960669d
MD5 90a118de594b637b971a329e7cbe40a6
BLAKE2b-256 accf000905a7f2957af51151494cf4ab80f1d45e3e3d861de446f28e58b430c4

See more details on using hashes here.

File details

Details for the file data_transformation_library_jmarci-0.1.3-py3-none-any.whl.

File metadata

File hashes

Hashes for data_transformation_library_jmarci-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 d452bf7b20b2550977c04454c103ee42211b3547172615679317d4c9614c1c35
MD5 d24b648d9cec059e2ef3e7af92cfe3a3
BLAKE2b-256 f6f94b185e3b2d46ebb8671585477b8f41d8bc56cd86f1a0434c8d2437cd8d5e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page