Skip to main content

Machine Learning tools for LETU FENNEC senior design team

Project description

fennec-ml

A suite of data science and machine learning tools for use by the LeTourneau University FENNEC senior design team.

Fennec Logo PyPI version

Instalation

pip install fennec-ml

Quick Start

Data Utils

import os
import fennec_ml as fn

# setup
root_dir = os.getcwd()
excel_dir = os.path.join(root_dir, "Raw_Data")
csv_dir = os.path.join(root_dir, "Proccessed_Data")
timesteps = 60

# excel to csv
fn.folder_cleaner(excel_dir)

# normalize and get labels
norm_data = fn.normalize(csv_dir)
labels = fn.get_CG_labels(csv_dir)

# segment and sort into train, validate, and test datasets
dataset_dict = fn.segment_and_split(norm_data, labels, timesteps)

# use dataset_dict
training_sets = dataset_dict['Training_Set']['sets']
training_labels = dataset_dict['Training_Set']['labels']

Project Structure

myproject/
├── data/
│   ├── raw_data/
│   │   ├── flight123_AA_L.xlxs
│   │   └── ...
│   └── proccessed_data/
│       ├── flight123_AA_L.csv
│       └── ...
├── saved_models/
│   └── ...
├── vars_of_interest.json
├── project_dev.ipynb
└── project_training.py

Features

data_cleaner()

Preprocesses .xlsx files into fennec question-usefull .csv files.

Args:

  • filepath (string): The .xlsx file to process.
  • savepath (string): The folder to save the .csv file.
  • overwrite (bool): Skips the overwrite checker if true.
  • skip (bool): Skips duplicate files instead of checking or overwriting if true.
  • varspath (string): The vars-of-interest.json path. Defaults to same folder as THIS script.

Relies on the vars_of_interest.json file to determine what data is wanted

fn.data_cleaner(excel_filename, overwrite= True)

folder_cleaner()

Preprocesses a folder of .xlsx files into fennec question-usefull .csv files.

Args:

  • excel_dir (string): The folder of .xlsx files to process.
  • savepath (string): The folder to save the .csv file.
  • overwrite (bool): Skips the overwrite checker if true.
  • skip (bool): Skips duplicate files instead of checking or overwriting if true.
  • varspath (string): The vars-of-interest.json path. Defaults to same folder as this script.

Relies on the vars_of_interest.json file to determine what data is wanted

fn.folder_cleaner(excel_dir, skip= True)

normalize()

Return a 3D array of normalized data from cleaned csv's
Note: Normalizing means scaling the data between the min and max values

Args:

  • csv_dir (string): The path (including the folder name) of cleaned data
  • weights (list): An optional list of weights corresponding to each column
  • offsets (list): An optional list of offsets corresponding to column

Returns:

  • norm_data (list): A list of numpy arrays holding normalized data
scaled_data = fn.normalize(csv_dir)

standardize()

Return a 3D array of standardized data from cleaned csv's
Note: Standarizing means scaling the data so the mean = 0 and the std deviation = 1

Args:

  • csv_dir (string): The path (including the folder name) of cleaned data
  • weights (list): An optional list of weights corresponding to each column
  • offsets (list): An optional list of offsets corresponding to column

Returns:

  • stand_data (list): A list of numpy arrays holding STANDARDIZED data
scaled_data = fn.standardize(csv_dir)

get_2D_CG_labels()

Reads all filenames in a folder and returns 2d CG characterization labels

Args:

  • csv_dir (string): Directory of .csv files from which to get labels

Returns:

  • labels (list): A list of all the characterization labels
cg_labels = fn.get_CG_labels(csv_dir)

get_1D_CG_labels()

Reads all filenames in a folder and returns 1d CG characterization labels

Args:

  • csv_dir (string): Directory of .csv files from which to get labels

Returns:

  • labels (list): A list of all the characterization labels
cg_labels = fn.get_CG_labels(csv_dir)

segment_and_split()

Segments, labels, and sorts data into dataset dictionary

Args:

  • input_data (list): List of numpy arrays, 1 per proccessed and scaled files
  • input_labels (list): List of characterization lables, 1 per file (should correspond to input_data)
  • timesteps (int): length of desired segments
  • train_split (float): Percentage of segments to save as training segments
  • validate_split (float): Percentage of segments to save as validation segments

Returns:

  • output (dict): 3 labels: "Training_Set", "Validation_Set", and "Testing_Set"
    Each set has the follwing labels: "sets" and "labels"
    • "sets" : list of sets, corresponds to "labels"
    • "labels" : list of labels, corresponds to "sets"
data_dict = fn.segment_and_split(all_data, all_labels, 60)

train_set = data_dict['Training_Set']['sets']
train_labels = data_dict['Training_Set']['labels']

Contributers

  • Luke Fagg (Team Lead)
  • Micah Yarbrough (Pilot ID)
  • Wills Kookogey (Fault ID)
  • Justin Hawk (3D CG)

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fennec_ml-0.5.0.tar.gz (9.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

fennec_ml-0.5.0-py3-none-any.whl (9.2 kB view details)

Uploaded Python 3

File details

Details for the file fennec_ml-0.5.0.tar.gz.

File metadata

  • Download URL: fennec_ml-0.5.0.tar.gz
  • Upload date:
  • Size: 9.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.11

File hashes

Hashes for fennec_ml-0.5.0.tar.gz
Algorithm Hash digest
SHA256 fe30d1bd43779876850b678e08f5763eed9e8949a44037e0f504e40d66fae082
MD5 bad7a6b3cf9303faa197e1532c8ab3fb
BLAKE2b-256 331d73ad37da82ad37d32fe3da7430076ed4499e4e604d6ceaa44f073a5a0e90

See more details on using hashes here.

File details

Details for the file fennec_ml-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: fennec_ml-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 9.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.11

File hashes

Hashes for fennec_ml-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 2119d3367970e444417c622fa18e16f3016b666aced46801a5f7c70839ca85ea
MD5 c4e925c7bf524627b264fcc7f74f51ac
BLAKE2b-256 576bff914e6e8659da4644feb3aa22cbe3e447b5fea93cd6ae01f5cab72ca685

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page