HollerithMLTrainTrack
is a Python library designed to track and analyze machine learning model training sessions, integrating key metrics such as training duration, feature analysis, and environmental impact metrics through CodeCarbon integration. It aims to provide developers and researchers with insights into their model's performance and its carbon footprint, enhancing transparency and accountability in ML projects.
Installation
Install HollerithMLTrainTrack with pip:
pip install hollerithmltraintrack
Quick Start
To get started with MLModelTracker, follow this simple example:
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.impute import SimpleImputer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import OneHotEncoder
from hollerithmltraintrack.main import MLModelTrackerInterface
# Load the data
titanic_url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
titanic_data = pd.read_csv(titanic_url)
X = titanic_data.drop(columns=['Survived'])
y = titanic_data['Survived']
# Split the data into training and test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Determine the column types
numeric_features = X_train.select_dtypes(include=['int64', 'float64']).columns
categorical_features = X_train.select_dtypes(include=['object']).columns
print("Number of numeric features:", len(numeric_features))
print("Number of categorical features:", len(categorical_features))
# Define the column transformations
preprocessor = ColumnTransformer(
transformers=[
('numeric', SimpleImputer(strategy='mean'), numeric_features),
('categorical', OneHotEncoder(), categorical_features)
],
remainder='passthrough'
)
# Process the training data
X_train_processed = preprocessor.fit_transform(X_train)
# Initialize the model
model = RandomForestClassifier(n_estimators=5, max_depth=2)
# Create the MLModelTrackerInterface instance
tracker_interface = MLModelTrackerInterface()
# Track the model training with the processed training data
with tracker_interface.track_training(model, X_train_processed, y_train, preprocessor, filename="custom_file_name.csv"):
model.fit(X_train_processed, y_train)
# Retrieve and print the tracked training information
tracked_info = tracker_interface.get_tracked_info()
print(tracked_info)
This example demonstrates how to track the training of a RandomForestClassifier on the Iris dataset. More detailed examples can be found in the examples folder.
Features
- Model Training Tracking: Capture detailed information about each training session, including model parameters, training duration, and feature counts.
- CodeCarbon Integration: Automatically track and report the carbon emissions associated with model training, leveraging the CodeCarbon library.
- CSV Export: Export collected data into a CSV file for further analysis or documentation purposes.
Advanced Usage
Docs to be added
Limitations and Known Issues
- Currently supports only sklearn models. Support for other frameworks is planned for future releases.
- The feature extraction only works when using the sklearn
ColumnTransformer - CodeCarbon integration relies on external services for emissions data, which may be subject to availability.
Contributing
Contributions are welcome! If you'd like to contribute, please:
- Fork the repository.
- Create a new branch for your feature.
- Submit a pull request.
Please refer to CONTRIBUTING.md for detailed contribution guidelines.
License
MLModelTracker is released under the MIT License. See the LICENSE file for more details.
Contact
For support or queries, please reach out to us at lukasgro63@gmail.com.
Feel free to report any issues or suggestions!
Metadata
Release files for HollerithMLTrainTrack 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| HollerithMLTrainTrack-0.1.0.tar.gz | 5.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| HollerithMLTrainTrack-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 13.0 kB
Release files / HollerithMLTrainTrack-0.1.0.tar.gz
| Download URL | HollerithMLTrainTrack-0.1.0.tar.gz |
|---|---|
| Size | 5.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
36ed48f9c2e35289a476e1431e1a748eddc6739d8e63063b82f79ef9cb0bcab2
|
|
BLAKE2b-256 checksum How to use checksums |
87c2be0363b1b3f2cd3f817cba8ad7c02e9210178fdfda1fceb3fad7e30d6a16
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.0.0 CPython/3.11.7
|
Release files / HollerithMLTrainTrack-0.1.0-py3-none-any.whl
| Download URL | HollerithMLTrainTrack-0.1.0-py3-none-any.whl |
|---|---|
| Size | 7.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
100aa15e56b6b2976d98d8ce3ceded1a4b242fddc947f3ea2c82c7e99e24ce6f
|
|
BLAKE2b-256 checksum How to use checksums |
bc2bf769d90800247f23085bdc7a0b643501def6a841c84b596015e75022d3e0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.0.0 CPython/3.11.7
|