Software Development ML Practice 1
Machine learning project focused on asteroid impact-risk analysis. The goal is to explore a dataset of near-Earth objects, analyze their risk-related features, train a baseline predictive model, and evaluate its performance through a structured machine learning pipeline.
This project follows a Cookiecutter Data Science-style organization and includes both an exploratory notebook and modular Python scripts for data preparation, feature engineering, visualization, model training, and prediction.
Documentation
The publish documentation is at: https://altamirano146080.github.io/asteroid_impact/
Overview
The repository analyzes a dataset related to potential asteroid impact risk. It includes:
- dataset exploration and profiling
- missing value analysis
- feature selection and target transformation
- visualization of risk-related patterns
- baseline neural network training
- model evaluation and predictions
The model uses asteroid characteristics such as encounter velocity, absolute magnitude, diameter, Palermo scale, and potential impact dates to predict the logarithm of the impact probability.
Data
This project uses the sentry-impact-risk dataset, which provides information on near-Earth objects and their potential collision risks.
- Source: The dataset is downloaded via the Hugging Face hub (juliensimon/sentry-impact-risk), curated by Julien Simon. The original observational data originates from the NASA/JPL Center for Near-Earth Object Studies (CNEOS) Sentry system.
- How and why it is used: The dataset is utilized to demonstrate a complete machine learning pipeline for regression tasks. We extract physical and kinematic features of the asteroids—such as encounter velocity, absolute magnitude, estimated diameter, and the Palermo scale—to train a baseline neural network. Because the actual impact probabilities are astronomically small, we transform the target variable to
log10(impact_probability)to enable stable model training and evaluation. - License: The dataset is provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Any redistribution or reuse of this data is subject to its applicable license and attribution requirements.
Project structure
.
├── LICENSE
├── Makefile
├── README.md
├── pyproject.toml
├── setup.cfg
├── data/
│ ├── external/
│ ├── interim/
│ ├── processed/
│ │ ├── dataset.csv
│ │ ├── features.csv
│ │ ├── labels.csv
│ │ └── predictions.csv
│ └── raw/
│ └── dataset.csv
├── docs/
├── models/
│ ├── feature_scaler.joblib
│ └── impact_probability_model.keras
├── notebooks/
│ └── explore_dataset.ipynb
├── references/
├── reports/
│ ├── model_metrics.csv
│ ├── training_history.csv
│ └── figures/
│ ├── correlation_heatmap.png
│ ├── impact_probability_distribution.png
│ ├── velocity_vs_palermo.png
│ ├── magnitude_vs_palermo.png
│ └── potential_impact_timeline.png
├── asteroid_impact/
│ ├── __init__.py
│ ├── config.py
│ ├── dataset.py
│ ├── features.py
│ ├── plots.py
│ └── modeling/
│ ├── __init__.py
│ ├── train.py
│ └── predict.py
└── uv.lock
Repository components
notebooks/explore_dataset.ipynb
Contains the complete exploratory analysis workflow, including:
- loading the dataset
- inspecting rows, columns, and data types
- checking missing values
- analyzing correlations and distributions
- visualizing asteroid risk properties
- preprocessing the data
- training and evaluating the baseline model
asteroid_impact/config.py
Defines the main project directories and creates the required folders for:
- raw data
- intermediate data
- processed data
- trained models
- reports and figures
asteroid_impact/dataset.py
Downloads or loads the asteroid impact-risk dataset and saves a local sample in the raw data directory. It also creates a cleaned dataset for the following pipeline steps.
asteroid_impact/features.py
Creates the input features and target variable by:
- removing incomplete rows
- applying a base-10 logarithmic transformation to
impact_probability - selecting numerical predictive variables
- saving
features.csvandlabels.csv
asteroid_impact/plots.py
Generates the exploratory data analysis visualizations, including:
- correlation heatmap
- impact probability distribution
- encounter velocity versus Palermo scale
- absolute magnitude versus Palermo scale
- potential impact timeline
asteroid_impact/modeling/train.py
Trains and evaluates the baseline neural network model. This script:
- splits the data into training and test sets
- standardizes the input features
- trains the TensorFlow model
- calculates MAE, RMSE, and R² metrics
- saves the model, scaler, metrics, and training history
asteroid_impact/modeling/predict.py
Loads the trained model and feature scaler, generates predictions for the processed features, and saves the results to a CSV file.
Requirements
This project uses Python 3.13. The required dependencies are specified in pyproject.toml and managed with uv.
To install the dependencies, run:
uv sync
To create the virtual environment with Python 3.13 and install the dependencies:
make setup
How to run
The project includes a Makefile with commands for each stage of the machine learning workflow.
Install dependencies
make requirements
Download and prepare the dataset
make data
Generate features and labels
make features
Generate exploratory plots
make plots
Train the model
make train
Generate predictions
make predict
Run the complete pipeline
make pipeline
The complete pipeline runs the following steps in order:
data → features → plots → train → predict
Open Jupyter Lab
make notebook
View all available commands
make help
Code quality
To check the source code with Flake8, isort, and Black:
make lint
To format the source code automatically:
make format
To remove Python cache files:
make clean
Generated outputs
The trained model and scaler are saved in:
models/impact_probability_model.kerasmodels/feature_scaler.joblib
The evaluation results are saved in:
reports/model_metrics.csvreports/training_history.csv
The exploratory plots are saved in:
reports/figures/
The predictions are saved in:
data/processed/predictions.csv
Notes
- The project uses a sample of up to 500 observations for the initial analysis.
- Rows with missing values are removed before model training.
- The target variable is transformed using
log10(impact_probability)because the original probabilities are very small. - The notebook is useful for interactive exploration and presentation.
- The Python modules provide a more maintainable and reusable version of the workflow.
License
This project is licensed under the MIT License. See the LICENSE file for details.
Authors
- Ruth Altamirano Trujillo
- Malena Flores Chacón
- Odei Martinez de Morentin
Metadata
Release files for asteroid-impact-sdoml 0.0.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| asteroid_impact_sdoml-0.0.3.tar.gz | 62.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| asteroid_impact_sdoml-0.0.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 128.4 kB
Release files / asteroid_impact_sdoml-0.0.3.tar.gz
| Download URL | asteroid_impact_sdoml-0.0.3.tar.gz |
|---|---|
| Size | 62.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6d5517aa5454286eb7b003c339ee9a3e0869c1aedb6a6f6156fe3ee1da9faeb4
|
|
BLAKE2b-256 checksum How to use checksums |
b51f442489615c188bbb0011f86c8f98649e9066112edf69cb92b4b579d2b21e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.22
|
Release files / asteroid_impact_sdoml-0.0.3-py3-none-any.whl
| Download URL | asteroid_impact_sdoml-0.0.3-py3-none-any.whl |
|---|---|
| Size | 65.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
76d971419c688090ef9346718e576db279ba1c4a3682bf2ea36925f0c1e5447b
|
|
BLAKE2b-256 checksum How to use checksums |
3069c14ca951ea349e9a60b4a549d5d37b733f3c0d327aed342673018473ca1d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.22
|