Skip to main content

Benchmark-Adv-ML

Benchmark-Adv-ML is a Python package designed to facilitate advanced benchmarking and analysis of machine learning models. It provides comprehensive pipelines for model stability evaluation, autoencoder training, and survival clustering analysis, enabling users to evaluate model performance, generate predictions, and visualize results through various plots, including AUC curves, feature importance charts, and Kaplan-Meier survival plots.

Table of Contents

Features

  • Model Stability Evaluation: Automatically runs multiple machine learning models (Logistic Regression, Support Vector Classifier, Random Forest Classifier) across multiple runs to assess stability and performance.
  • Autoencoder Training: Implements an autoencoder for dimensionality reduction and feature extraction, customizable with various hyperparameters.
  • Survival Clustering Analysis: Performs clustering on patient features and integrates clinical data to generate Kaplan-Meier survival plots and log-rank tests.
  • Prediction and Metrics Generation: Generates and saves predictions, feature importance scores, and various performance metrics for each model and run.
  • Aggregation of Results: Aggregates results across runs and models for comprehensive analysis, facilitating comparison and evaluation.
  • Visualization Tools: Generates plots including AUC curves, AUC box plots, feature importance charts, radar charts for model performance comparison, and survival analysis plots.

Installation

You can install the package directly from PyPI:

pip install benchmark-adv-ml

Alternatively, install from source:

git clone https://github.com/yourusername/benchmark-adv-ml.git
cd benchmark-adv-ml
pip install .

Useage

The package provides a command-line interface (CLI) for ease of use. Below are examples of how to use each component.

Download example data

wget https://github.com/VatsalPatel18/benchmark-adv-ml/blob/master/temp_data.csv

Benchmark Machine Learning Models

Run the benchmark ML pipeline to evaluate model stability across multiple runs.

benchmark-adv-ml benchmark --data ./your_dataset.csv --output ./final_results --prelim_output ./prelim_results --n_runs 10 --seed 42

Train Autoencoder Model

Train and evaluate an autoencoder model for feature extraction.

benchmark-adv-ml autoencoder --data ./your_dataset.csv --sampleID 'PatientID' --output_dir ./final_results --prelim_output ./prelim_results --latent_dim 10 --epochs 50 --batch_size 32 --validation_split 0.1 --test_size 0.2 --seed 42

Survival Clustering Analysis

benchmark-adv-ml survival_clustering --data_path ./latent_features.csv --clinical_df_path ./clinical_data.csv --save_dir ./final_results

Command-Line Arguments

Common Arguments

  • --data: Path to the existing CSV file containing the dataset.
  • --output: Directory to save the final results and plots.
  • --prelim_output: Directory to save the preliminary results (predictions).
  • --seed: Seed for random state (default is 42).

Benchmark Command Arguments

  • --target: Target column name in the dataset (default: 'label').
  • --n_runs: Number of runs for model stability evaluation (default: 20).

Autoencoder Command Arguments

  • --sampleID: Column name representing the sample or patient ID (default: 'sampleID').
  • --latent_dim: Dimensionality of the latent space (default: input_dim // 8).
  • --epochs: Number of training epochs (default: 50).
  • --batch_size: Training batch size (default: 32).
  • --validation_split: Proportion of training data to use as validation set (default: 0.1).
  • --test_size: Proportion of data to use as test set (default: 0.2).
  • --early_stopping: Enable early stopping (use flag to activate).
  • --patience: Patience for early stopping (default: 5).
  • --checkpoint: Enable model checkpointing (use flag to activate).

Survival Clustering Command Arguments

  • --data_path: Path to the CSV file containing patient features.
  • --clinical_df_path: Path to the CSV file containing clinical data.
  • --save_dir: Directory to save the results.

Dependencies

  • Python 3.11+
  • numpy
  • pandas
  • scikit-learn
  • matplotlib
  • seaborn
  • tensorflow
  • lifelines
  • yellowbrick

License

This project is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. See the LICENSE file for details.

Author

Vatsal Patel - VatsalPatel18

Metadata

Release files for benchmark-adv-ml 0.2.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for benchmark-adv-ml 0.2.8
File Size Uploaded
benchmark_adv_ml-0.2.8.tar.gz 80.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for benchmark-adv-ml 0.2.8
File Interpreter ABI Platform
benchmark_adv_ml-0.2.8-py3-none-any.whl Python 3 none any Details

Total release size: 169.5 kB

Release files / benchmark_adv_ml-0.2.8.tar.gz

Download URL benchmark_adv_ml-0.2.8.tar.gz
Size 80.2 kB
Tags Source
SHA-256 checksum
How to use checksums
80226b9c96a8e8b9892689018f80ff70fb4b32e39f3e7a9cb3220416718deccc
BLAKE2b-256 checksum
How to use checksums
db8e563041512650fe4608760169fa127afdcca454c911cfe10b0c27b1604336
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.3 CPython/3.11.7 Linux/6.8.0-45-generic

Release files / benchmark_adv_ml-0.2.8-py3-none-any.whl

Download URL benchmark_adv_ml-0.2.8-py3-none-any.whl
Size 89.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
af0e7b3253cbec0252c54dd36b29a84b30712e406b24ea79f73793e1815b9973
BLAKE2b-256 checksum
How to use checksums
4e75fe37b2d6cb9fd3177fbd67cc3bf25ed5c1685e8ecadb532aed07984fb467
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.3 CPython/3.11.7 Linux/6.8.0-45-generic

Release history Release notifications | RSS feed

This release

0.2.8 This release

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.2

2 release files

0.2.0

2 release files

0.1.24

2 release files

0.1.23

2 release files

0.1.22

2 release files

0.1.21

2 release files

0.1.20

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

0.1.16

2 release files

0.1.15

2 release files

0.1.14

2 release files

0.1.13

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.4

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page