Skip to main content

A Python library to find optimal configurations for regressors

Project description

OptimalRegressors

OptimalRegressors is a Python library designed to find the optimal configurations for decision tree and random forest regressors. It automates hyperparameter tuning, such as determining the optimal number of leaf nodes, using validation data to improve model performance.

Features

  • Automatically optimizes hyperparameters for:
    • DecisionTreeRegressor
    • RandomForestRegressor
  • Allows custom candidate values for hyperparameter tuning.
  • Optionally fits the optimized model for direct use.
  • Simple, lightweight, and easy to use.

Installation

To install OptimalRegressors, follow the steps below:

Clone the Repository

git clone https://github.com/RehanTaneja/OptimalRegressors.git
cd OptimalRegressors

Install Dependencies

Use pip to install the required dependencies:

pip install -r requirements.txt

Usage

regressors.py

OptimalDecisionTreeRegressor

  • Parameters:
    • trainX,trainy Training data (features & labels)
    • valX,valy Validation data (features & labels)
    • candidate_nodes (list, optional): A list of max_leaf_nodes values to test. Default: [5, 50, 500, 5000]
    • candidate_splits (list, optional): A list of min_sample_split values to test. Default: [2,3,4,5,6,7,8,9,10]
    • candidate_leaves (list,optional): A list of min_sample_leaf values to test. Default: [1,2,3,4,5]
    • candidate_depths (list,optional): A list of max_depth values to test. Default: [3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20]
    • fit (bool, optional): Whether to fit the returned model on the training data. Default: False
  • Returns:
    • model: A DecisionTreeRegressor instance with the optimal configuration

Find the best max_leaf_nodes for a decision tree:

from OptimalRegressor import OptimalDecisionTreeRegressor

# Example data (replace with your dataset)
trainX, trainy = [[1], [2], [3]], [2.5, 3.5, 5.0]
valX, valy = [[1.5], [2.5]], [3.0, 4.0]

# Get the optimal DecisionTreeRegressor
model, nodes = OptimalDecisionTreeRegressor(trainX, trainy, valX, valy)

print("Optimal max_leaf_nodes:", nodes)

OptimalRandomForestRegressor

  • Parameters:
    • trainX, trainy: Training data (features and labels)
    • vsalX, valy: Validation data (features and labels)
    • candidate_nodes (list, optional): A list of max_leaf_nodes values to test. Default: [5, 50, 500, 5000]
    • candidate_estimators (list,optional): A list of n_estimator values to test. Default: [50,100,200,300,400,500]
    • fit (bool, optional): Whether to fit the returned model on the training data. Default: False
  • Returns:
    • model: A RandomForestRegressor instance with the optimal configuration

Find the best max_leaf_nodes value for a random forest regressor

from OptimalRegressor import OptimalRandomForestRegressor

# Example data (replace with your dataset)
trainX, trainy = [[1], [2], [3]], [2.5, 3.5, 5.0]
valX, valy = [[1.5], [2.5]], [3.0, 4.0]

# Get the optimal RandomForestRegressor
model, nodes = OptimalRandomForestRegressor(trainX, trainy, valX, valy)

print("Optimal max_leaf_nodes:", nodes)

OptimalDecisionTreeRegressors.py

OptimalMaxLeafNodes

  • Parameters:
  • candidate_nodes: A list of max_leaf_nodes values to test
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for max_leaf_node with minimum minimum absolute error

OptimalMinSampleSplit

  • Parameters:
  • candidate_splits: A list of min_sample_split values to test
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for min_sample_split with minimum minimum absolute error

OptimalMinSampleLeaf

  • Parameters:
  • candidate_nleaves: A list of min_sample_leaf values to test
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for min_sample_leaf with minimum minimum absolute error

OptimalMaxDepth

  • Parameters:
  • candidate_nodes: A list of max_depth values to test
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for max_depth with minimum minimum absolute error

OptimalRandomForestRegressors.py

OptimalMaxLeafNodes

  • Parameters:
  • candidate_nodes: A list of max_leaf_nodes values to test
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for max_leaf_node with minimum minimum absolute error

OptimalNEstimators

  • Parameters:
  • candidate_estimators: A list of n_estimators values to test
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for n_estimators with minimum minimum absolute error

OptimalBootstrap

  • Parameters:
  • trainX, trainy: Training data (features and labels)
  • valX, valy: Validation data (features and labels)
  • mae: Original Mean Absolute Error
  • Returns:
    • Optimal value for bootstrap with minimum minimum absolute error

Repository Structure

OptimalRegressors/ │ ├── OptimalDecisionTreeRegressors.py # Methods for getting optimal hyperparameters for DecisionTree |-- OptimalRandomForestRegressors.py # Methods for getting optimal hyperparameters for RandomForest |-- regressor.py # Main library file ├── requirements.txt # Python dependencies ├── LICENSE # License information └── README.md # Project documentation

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

optimal_regressors-0.8.0.tar.gz (3.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

optimal_regressors-0.8.0-py3-none-any.whl (4.7 kB view details)

Uploaded Python 3

File details

Details for the file optimal_regressors-0.8.0.tar.gz.

File metadata

  • Download URL: optimal_regressors-0.8.0.tar.gz
  • Upload date:
  • Size: 3.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.12.6

File hashes

Hashes for optimal_regressors-0.8.0.tar.gz
Algorithm Hash digest
SHA256 76ec9243e9076e84461b00377a612d7f68e1c8f21f65ce4b945251b49a75b461
MD5 3c7c6d21b3f8ff5af2f4f07d255462e3
BLAKE2b-256 6fa758890a080eb13bc4ab2b8548b71ffb0def0577a26984b95fb5166a3e8542

See more details on using hashes here.

File details

Details for the file optimal_regressors-0.8.0-py3-none-any.whl.

File metadata

File hashes

Hashes for optimal_regressors-0.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 5045f1116643a75045ebfd1e8a462fefeb0e930bc8dfc166bd649cd6f726f235
MD5 82518670371c0d7eafe01c4d241581ac
BLAKE2b-256 43b11fced7fd0e136003629f50f262a27b2c16a70832dc53f7cb111767dc2bbe

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page