Skip to main content

Software Development ML Practice 1

Machine learning project focused on asteroid impact-risk analysis. The goal is to explore a dataset of near-Earth objects, analyze their risk-related features, train a baseline predictive model, and evaluate its performance through a structured machine learning pipeline.

This project follows a Cookiecutter Data Science-style organization and includes both an exploratory notebook and modular Python scripts for data preparation, feature engineering, visualization, model training, and prediction.

Overview

The repository analyzes a dataset related to potential asteroid impact risk. It includes:

  • dataset exploration and profiling
  • missing value analysis
  • feature selection and target transformation
  • visualization of risk-related patterns
  • baseline neural network training
  • model evaluation and predictions

The model uses asteroid characteristics such as encounter velocity, absolute magnitude, diameter, Palermo scale, and potential impact dates to predict the logarithm of the impact probability.

Project structure

.
├── LICENSE
├── Makefile
├── README.md
├── pyproject.toml
├── setup.cfg
├── data/
│   ├── external/
│   ├── interim/
│   ├── processed/
│   │   ├── dataset.csv
│   │   ├── features.csv
│   │   ├── labels.csv
│   │   └── predictions.csv
│   └── raw/
│       └── dataset.csv
├── docs/
├── models/
│   ├── feature_scaler.joblib
│   └── impact_probability_model.keras
├── notebooks/
│   └── explore_dataset.ipynb
├── references/
├── reports/
│   ├── model_metrics.csv
│   ├── training_history.csv
│   └── figures/
│       ├── correlation_heatmap.png
│       ├── impact_probability_distribution.png
│       ├── velocity_vs_palermo.png
│       ├── magnitude_vs_palermo.png
│       └── potential_impact_timeline.png
├── software_development_ml_practice_1/
│   ├── __init__.py
│   ├── config.py
│   ├── dataset.py
│   ├── features.py
│   ├── plots.py
│   └── modeling/
│       ├── __init__.py
│       ├── train.py
│       └── predict.py
└── uv.lock

Repository components

notebooks/explore_dataset.ipynb

Contains the complete exploratory analysis workflow, including:

  • loading the dataset
  • inspecting rows, columns, and data types
  • checking missing values
  • analyzing correlations and distributions
  • visualizing asteroid risk properties
  • preprocessing the data
  • training and evaluating the baseline model

software_development_ml_practice_1/config.py

Defines the main project directories and creates the required folders for:

  • raw data
  • intermediate data
  • processed data
  • trained models
  • reports and figures

software_development_ml_practice_1/dataset.py

Downloads or loads the asteroid impact-risk dataset and saves a local sample in the raw data directory. It also creates a cleaned dataset for the following pipeline steps.

software_development_ml_practice_1/features.py

Creates the input features and target variable by:

  • removing incomplete rows
  • applying a base-10 logarithmic transformation to impact_probability
  • selecting numerical predictive variables
  • saving features.csv and labels.csv

software_development_ml_practice_1/plots.py

Generates the exploratory data analysis visualizations, including:

  • correlation heatmap
  • impact probability distribution
  • encounter velocity versus Palermo scale
  • absolute magnitude versus Palermo scale
  • potential impact timeline

software_development_ml_practice_1/modeling/train.py

Trains and evaluates the baseline neural network model. This script:

  • splits the data into training and test sets
  • standardizes the input features
  • trains the TensorFlow model
  • calculates MAE, RMSE, and R² metrics
  • saves the model, scaler, metrics, and training history

software_development_ml_practice_1/modeling/predict.py

Loads the trained model and feature scaler, generates predictions for the processed features, and saves the results to a CSV file.

Requirements

This project uses Python 3.13. The required dependencies are specified in pyproject.toml and managed with uv.

To install the dependencies, run:

uv sync

To create the virtual environment with Python 3.13 and install the dependencies:

make setup

How to run

The project includes a Makefile with commands for each stage of the machine learning workflow.

Install dependencies

make requirements

Download and prepare the dataset

make data

Generate features and labels

make features

Generate exploratory plots

make plots

Train the model

make train

Generate predictions

make predict

Run the complete pipeline

make pipeline

The complete pipeline runs the following steps in order:

data → features → plots → train → predict

Open Jupyter Lab

make notebook

View all available commands

make help

Code quality

To check the source code with Flake8, isort, and Black:

make lint

To format the source code automatically:

make format

To remove Python cache files:

make clean

Generated outputs

The trained model and scaler are saved in:

  • models/impact_probability_model.keras
  • models/feature_scaler.joblib

The evaluation results are saved in:

  • reports/model_metrics.csv
  • reports/training_history.csv

The exploratory plots are saved in:

  • reports/figures/

The predictions are saved in:

  • data/processed/predictions.csv

Notes

  • The project uses a sample of up to 500 observations for the initial analysis.
  • Rows with missing values are removed before model training.
  • The target variable is transformed using log10(impact_probability) because the original probabilities are very small.
  • The notebook is useful for interactive exploration and presentation.
  • The Python modules provide a more maintainable and reusable version of the workflow.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Authors

  • Ruth Altamirano Trujillo
  • Malena Flores Chacón
  • Odei Martinez de Morentin

Metadata

Release files for software_development_ml_practice_1 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for software_development_ml_practice_1 0.0.1
File Size Uploaded
software_development_ml_practice_1-0.0.1.tar.gz 17.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for software_development_ml_practice_1 0.0.1
File Interpreter ABI Platform
software_development_ml_practice_1-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 40.9 kB

Release files / software_development_ml_practice_1-0.0.1.tar.gz

Download URL software_development_ml_practice_1-0.0.1.tar.gz
Size 17.5 kB
Tags Source
SHA-256 checksum
How to use checksums
7623a8a4f618dc5eb534f77c27cbc948568b565446e2b10e89c26deac1ae91f7
BLAKE2b-256 checksum
How to use checksums
a7a40160a16ee693d71d359be03045e6228695de3fdb33c3cfae5335ec9f1a17
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.22

Release files / software_development_ml_practice_1-0.0.1-py3-none-any.whl

Download URL software_development_ml_practice_1-0.0.1-py3-none-any.whl
Size 23.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
eea065fec507edee54e76f8a481e440f5959225b8b313fa8deb2446353e04585
BLAKE2b-256 checksum
How to use checksums
805d66a8ead5009470c00ce719f29928da312365aad9ea66cc762eeb2bb43125
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.22

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page