Skip to main content

SnapBoost

PyPI version License: MIT scikit-learn

Heterogeneous Newton Boosting Machine (HNBM) — a gradient boosting framework that mixes decision trees and kernel ridge regressors instead of trees alone. The core HNBM framework is provided by the hnbm package; SnapBoost is a concrete implementation built on top of it.

Unlike XGBoost and LightGBM, which rely exclusively on decision trees as base learners, SnapBoost stochastically selects from a heterogeneous pool of learners at each boosting iteration. This lets the model capture both local, axis-aligned structure (trees) and smooth, global patterns (RBF kernel ridge).

This package is a Python/scikit-learn reimplementation inspired by SnapBoost: A Heterogeneous Boosting Machine (Parnell et al., NeurIPS 2020). See REFERENCES.md for papers, related work, and citation details.


Table of Contents


Features

Tag Description
gradient-boosting Second-order Newton boosting with gradient and Hessian weighting
heterogeneous-learners Mixes decision trees and kernel ridge regressors in one ensemble
classification Binary classification with logistic loss
regression Continuous targets with mean squared error loss
scikit-learn Implements the scikit-learn estimator API (fit, predict, score, …)
randomized-ensemble Stochastic base-learner selection per iteration

Installation

From PyPI (recommended):

pip install snapboost

From source:

git clone https://github.com/qiancapital/snapboost.git
cd snapboost
pip install .

Requirements: Python ≥ 3.8, NumPy, scikit-learn, tqdm, hnbm ≥ 0.2.2.


Quick Start

Classification

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from snapboost import SnapBoostClassifier

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

model = SnapBoostClassifier(
    num_iterations=100,
    learning_rate=0.1,
    random_state=42,
)
model.fit(X_train, y_train)

print("Accuracy:", model.score(X_test, y_test))
print("Probabilities shape:", model.predict_proba(X_test).shape)  # (n_samples, 2)
model.evaluate(X_test, y_test)  # prints log loss

Regression

from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from snapboost import SnapBoostRegressor

X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

model = SnapBoostRegressor(
    num_iterations=100,
    learning_rate=0.1,
    random_state=42,
)
model.fit(X_train, y_train)

print("R²:", model.score(X_test, y_test))
model.evaluate(X_test, y_test)  # prints RMSE

Examples & Results

Interactive Jupyter notebooks in static/ walk through classification, regression, and hyperparameter exploration. Each notebook trains SnapBoost and compares it against XGBoost and LightGBM on the same splits.

Notebook Dataset SnapBoost XGBoost LightGBM
Classification.ipynb Breast Cancer Wisconsin 97.2% accuracy 95.8% 96.5%
Regression.ipynb Diabetes R² 0.44, RMSE 55.7 R² 0.38, RMSE 58.4 R² 0.40, RMSE 57.7
Parameter_Exploration.ipynb Synthetic (piecewise + smooth) R² 0.986, RMSE 0.170 R² 0.986, RMSE 0.174 R² 0.987, RMSE 0.167

Run the notebooks locally:

pip install ".[examples]"
jupyter notebook static/

Classification

On the Breast Cancer dataset (250 boosting rounds), SnapBoost achieves the highest test accuracy and fewest misclassifications among the three boosters:

Test accuracy and error count vs XGBoost and LightGBM

Confusion matrix for SnapBoost on the held-out test set:

SnapBoost classification confusion matrix

Regression

On the Diabetes dataset (100 boosting rounds), SnapBoost improves R² and RMSE over tree-only baselines:

R², RMSE, and MAE comparison on Diabetes dataset

Predicted vs. actual disease progression on the test set:

Predicted vs actual scatter plot

SnapBoost fitted curve along BMI (other features held at training medians):

BMI vs target with SnapBoost fit

Residual distribution:

Regression residual histogram

Parameter exploration

On a synthetic dataset mixing piecewise-linear and sinusoidal structure, the notebook sweeps p_tree, tree depth ranges, and kernel ridge parameters. A mixed ensemble (p_tree=0.8) outperforms trees-only (p_tree=1.0, RMSE 0.174) and ridge-only (p_tree=0.0, RMSE 0.366):

Learned functions along one axis for different p_tree values

See Parameter_Exploration.ipynb for the full sweeps and baseline comparison tables.


API Reference

SnapBoostClassifier / SnapBoostRegressor

The recommended entry points (similar to XGBClassifier / XGBRegressor). A concrete HNBM that builds an ensemble from:

  • Decision trees with depths sampled uniformly from [min_max_depth, max_max_depth]
  • One RFF ridge regressor for smooth global fits

At each iteration, a learner is chosen with probability p_tree for trees (split evenly across depths) and 1 - p_tree for the ridge model.

from snapboost import SnapBoostClassifier, SnapBoostRegressor

clf = SnapBoostClassifier(
    num_iterations=100,
    learning_rate=0.1,
    p_tree=0.8,
    min_max_depth=4,
    max_max_depth=8,
    alpha=1.0,
    gamma=1.0,
    random_state=42,
    verbose=True,
)
clf.fit(X, y)

reg = SnapBoostRegressor(num_iterations=100, random_state=42)
reg.fit(X, y)

Methods

Method Classifier Regressor Description
fit(X, y) Train the ensemble
predict(X) Original class labels or continuous values
predict_proba(X) Class probabilities, shape (n_samples, 2)
decision_function(X) Raw logits
score(X, y) Accuracy or R²
evaluate(X, y) Prints and returns log loss or RMSE

SnapBoost

Legacy class that accepts a mode parameter ("classification" or "regression"). Prefer SnapBoostClassifier or SnapBoostRegressor for new code.

from snapboost import SnapBoost

model = SnapBoost(
    num_iterations=100,
    learning_rate=0.1,
    p_tree=0.8,
    min_max_depth=4,
    max_max_depth=8,
    alpha=1.0,
    gamma=1.0,
    mode="classification",  # or "regression"
    random_state=42,
    verbose=True,
)
model.fit(X, y)

Exact kernel ridge variant

For smaller datasets where an exact RBF kernel is preferable to random Fourier features, task-specific exact-kernel estimators are also available:

from snapboost import (
    SnapBoostKernelRidgeClassifier,
    SnapBoostKernelRidgeRegressor,
)

clf = SnapBoostKernelRidgeClassifier(random_state=42)
reg = SnapBoostKernelRidgeRegressor(random_state=42)

Exact kernel ridge has substantially higher memory and runtime costs than the default RFF learner. The old SnapBoost_KernelRidge name remains available for backward compatibility, but new code should use the task-specific classes.

HNBM

The abstract base class for building custom heterogeneous ensembles. Provided by the hnbm package — subclass or configure base_learners_ and probabilities_ before calling fit:

from sklearn.tree import DecisionTreeRegressor
from hnbm import HNBMClassifier, HNBMRegressor

class MyClassifier(HNBMClassifier):
    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        self.base_learners_ = [DecisionTreeRegressor(max_depth=5)]
        self.probabilities_ = [1.0]

Parameters

Shared (HNBM / SnapBoostClassifier / SnapBoostRegressor)

Parameter Type Default Description
num_iterations int 100 Number of boosting rounds
learning_rate float 0.1 Shrinkage applied to each learner's contribution
random_state non-negative int or None None Seed for learner selection and independently derived base-learner seeds
verbose bool False Show a tqdm progress bar during training

The legacy SnapBoost class also accepts a mode parameter ("classification" or "regression").

SnapBoost-specific

Parameter Type Default Description
p_tree float 0.9 Probability of selecting a decision tree (vs. ridge)
min_max_depth int 2 Minimum max_depth for trees in the pool
max_max_depth int 4 Maximum max_depth for trees in the pool
min_samples_leaf int 10 Minimum number of samples required in each decision-tree leaf
alpha float 1.0 L2 regularization for the RFF ridge regressor
gamma float 1.0 RBF kernel coefficient for random Fourier features
n_components int 100 Number of random Fourier features

Label conventions (classification): accepts any two distinct class labels. Predictions use the original labels, and probability columns follow classes_ order.


Docker

Build and run a container with SnapBoost pre-installed:

docker build -t snapboost .
docker run --rm snapboost

The default command verifies the import:

SnapBoost ready

Development

For normal development against the released HNBM dependency, install SnapBoost in editable mode and run the complete validation suite:

git clone https://github.com/qiancapital/snapboost.git
cd snapboost
python -m pip install -e ".[test]"
python -m pytest -q
python -m compileall -q snapboost tests

The pytest command must finish with all tests passing. To validate SnapBoost against a local sibling checkout of HNBM, install that checkout first:

python -m pip install -e ../hnbm
python -m pip install -e ".[test]"
python -m pytest -q

Run an individual test module or test while developing with:

python -m pytest -q tests/test_snapboost.py
python -m pytest -q tests/test_rff_learner.py
python -m pytest -q tests/test_snapboost.py::test_classifier_preserves_string_labels

The example notebooks require the separate examples dependencies:

python -m pip install -e ".[examples,test]"
jupyter notebook static/

CI runs the full test suite on every push and pull request, and again before a release distribution is built.

Releases are published to PyPI via GitHub Actions when a GitHub release is created.


References & Citation

If you use this package or the HNBM framework in research, please cite the original SnapBoost paper:

Thomas Parnell, Andreea Anghel, Małgorzata Łazuka, Nikolas Ioannou, Sebastian Kurella, Peshal Agarwal, Nikolaos Papandreou, and Haralampos Pozidis. SnapBoost: A Heterogeneous Boosting Machine. Advances in Neural Information Processing Systems, 33, 2020.

@inproceedings{parnell2020snapboost,
  title     = {{SnapBoost}: A Heterogeneous Boosting Machine},
  author    = {Parnell, Thomas and Anghel, Andreea and {\L}azuka, Ma{\l}gorzata and Ioannou, Nikolas and Kurella, Sebastian and Agarwal, Peshal and Papandreou, Nikolaos and Pozidis, Haralampos},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {33},
  pages     = {20872--20883},
  year      = {2020},
  eprint    = {2006.09745},
  doi       = {10.48550/arXiv.2006.09745}
}

Links: arXiv:2006.09745 · NeurIPS proceedings · IBM Research

For the full bibliography, related heterogeneous-boosting literature (KTBoost, DeepBoost, etc.), and notes on how this repo relates to the original IBM Snap ML implementation, see REFERENCES.md. Additional BibTeX entries are in CITATION.bib.


License

MIT — See LICENSE for full text.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

snapboost-0.1.7.tar.gz (22.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

snapboost-0.1.7-py3-none-any.whl (11.5 kB view details)

Uploaded Python 3

File details

Details for the file snapboost-0.1.7.tar.gz.

File metadata

  • Download URL: snapboost-0.1.7.tar.gz
  • Upload date:
  • Size: 22.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for snapboost-0.1.7.tar.gz
Algorithm Hash digest
SHA256 1e6b33c9a7ec43dd47055037103d0a4c4d3a6e45818cef78657cb5e8ed8712b7
MD5 3f0f9a7265a99be4439e70a67b6e2638
BLAKE2b-256 5440e283e73525570a8b473f9015c4d697fbcf09178a2392f6263c69255b7b72

See more details on using hashes here.

Provenance

The following attestation bundles were made for snapboost-0.1.7.tar.gz:

Publisher: python-publish.yml on QianCapital/snapboost

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file snapboost-0.1.7-py3-none-any.whl.

File metadata

  • Download URL: snapboost-0.1.7-py3-none-any.whl
  • Upload date:
  • Size: 11.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for snapboost-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 012262099f3bcf6e31e91fecb44134dfed165befc0245c45b5152f4c084dcd70
MD5 9ab6e1af40ad7d00d1321fcc0ca715d7
BLAKE2b-256 c50f92d18158905f9772c6eb1fd67ce9cf66b4a14c4a4543b6063bf44ceb5308

See more details on using hashes here.

Provenance

The following attestation bundles were made for snapboost-0.1.7-py3-none-any.whl:

Publisher: python-publish.yml on QianCapital/snapboost

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.2.0

2 files

1.0.0

2 files

0.2.1

2 files

0.2.0

2 files

This release

0.1.7 This release

2 files

0.1.6

2 files

0.1.5

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page