SnapBoost
Heterogeneous Newton Boosting Machine (HNBM) — a gradient boosting framework that mixes decision trees and kernel ridge regressors instead of trees alone. The core HNBM framework is provided by the hnbm package; SnapBoost is a concrete implementation built on top of it.
Unlike XGBoost and LightGBM, which rely exclusively on decision trees as base learners, SnapBoost stochastically selects from a heterogeneous pool of learners at each boosting iteration. This lets the model capture both local, axis-aligned structure (trees) and smooth, global patterns (RBF kernel ridge).
This package is a Python/scikit-learn reimplementation inspired by SnapBoost: A Heterogeneous Boosting Machine (Parnell et al., NeurIPS 2020). See REFERENCES.md for papers, related work, and citation details.
Table of Contents
- Features
- Installation
- Quick Start
- API Reference
- Parameters
- Docker
- Development
- References & Citation
- License
Features
| Tag | Description |
|---|---|
gradient-boosting |
Second-order Newton boosting with gradient and Hessian weighting |
heterogeneous-learners |
Mixes decision trees and kernel ridge regressors in one ensemble |
classification |
Binary classification with logistic loss |
regression |
Continuous targets with mean squared error loss |
scikit-learn |
Implements the scikit-learn estimator API (fit, predict, score, …) |
randomized-ensemble |
Stochastic base-learner selection per iteration |
Installation
From PyPI (recommended):
pip install snapboost
From source:
git clone https://github.com/qiancapital/snapboost.git
cd snapboost
pip install .
Requirements: Python ≥ 3.8, NumPy, scikit-learn, tqdm, hnbm ≥ 0.1.1.
Quick Start
Classification
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from snapboost import SnapBoostClassifier
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
model = SnapBoostClassifier(
num_iterations=100,
learning_rate=0.1,
random_state=42,
)
model.fit(X_train, y_train)
print("Accuracy:", model.score(X_test, y_test))
print("Probabilities shape:", model.predict_proba(X_test).shape) # (n_samples, 2)
model.evaluate(X_test, y_test) # prints log loss
Regression
from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from snapboost import SnapBoostRegressor
X, y = load_diabetes(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
model = SnapBoostRegressor(
num_iterations=100,
learning_rate=0.1,
random_state=42,
)
model.fit(X_train, y_train)
print("R²:", model.score(X_test, y_test))
model.evaluate(X_test, y_test) # prints RMSE
API Reference
SnapBoostClassifier / SnapBoostRegressor
The recommended entry points (similar to XGBClassifier / XGBRegressor). A concrete HNBM that builds an ensemble from:
- Decision trees with depths sampled uniformly from
[min_max_depth, max_max_depth] - One RFF ridge regressor for smooth global fits
At each iteration, a learner is chosen with probability p_tree for trees (split evenly across depths) and 1 - p_tree for the ridge model.
from snapboost import SnapBoostClassifier, SnapBoostRegressor
clf = SnapBoostClassifier(
num_iterations=100,
learning_rate=0.1,
p_tree=0.8,
min_max_depth=4,
max_max_depth=8,
alpha=1.0,
gamma=1.0,
random_state=42,
verbose=True,
)
clf.fit(X, y)
reg = SnapBoostRegressor(num_iterations=100, random_state=42)
reg.fit(X, y)
Methods
| Method | Classifier | Regressor | Description |
|---|---|---|---|
fit(X, y) |
✓ | ✓ | Train the ensemble |
predict(X) |
✓ | ✓ | Class labels (0/1) or continuous values |
predict_proba(X) |
✓ | Class probabilities, shape (n_samples, 2) |
|
decision_function(X) |
✓ | Raw logits | |
score(X, y) |
✓ | ✓ | Accuracy or R² |
evaluate(X, y) |
✓ | ✓ | Prints and returns log loss or RMSE |
SnapBoost
Legacy class that accepts a mode parameter ("classification" or "regression"). Prefer SnapBoostClassifier or SnapBoostRegressor for new code.
from snapboost import SnapBoost
model = SnapBoost(
num_iterations=100,
learning_rate=0.1,
p_tree=0.8,
min_max_depth=4,
max_max_depth=8,
alpha=1.0,
gamma=1.0,
mode="classification", # or "regression"
random_state=42,
verbose=True,
)
model.fit(X, y)
HNBM
The abstract base class for building custom heterogeneous ensembles. Provided by the hnbm package — subclass or configure base_learners_ and probabilities_ before calling fit:
from sklearn.tree import DecisionTreeRegressor
from hnbm import HNBMClassifier, HNBMRegressor
class MyClassifier(HNBMClassifier):
def __init__(self, **kwargs):
super().__init__(**kwargs)
self.base_learners_ = [DecisionTreeRegressor(max_depth=5)]
self.probabilities_ = [1.0]
Parameters
Shared (HNBM / SnapBoostClassifier / SnapBoostRegressor)
| Parameter | Type | Default | Description |
|---|---|---|---|
num_iterations |
int |
100 |
Number of boosting rounds |
learning_rate |
float |
0.1 |
Shrinkage applied to each learner's contribution |
random_state |
int or None |
None |
Seed for learner selection and tree fitting |
verbose |
bool |
True |
Show a tqdm progress bar during training |
The legacy SnapBoost class also accepts a mode parameter ("classification" or "regression").
SnapBoost-specific
| Parameter | Type | Default | Description |
|---|---|---|---|
p_tree |
float |
0.8 |
Probability of selecting a decision tree (vs. ridge) |
min_max_depth |
int |
4 |
Minimum max_depth for trees in the pool |
max_max_depth |
int |
8 |
Maximum max_depth for trees in the pool |
alpha |
float |
1.0 |
L2 regularization for the RFF ridge regressor |
gamma |
float |
1.0 |
RBF kernel coefficient for random Fourier features |
n_components |
int |
100 |
Number of random Fourier features |
Label conventions (classification): accepts 0/1 or -1/+1. Predictions are returned as 0/1.
Docker
Build and run a container with SnapBoost pre-installed:
docker build -t snapboost .
docker run --rm snapboost
The default command verifies the import:
SnapBoost ready
Development
git clone https://github.com/qiancapital/snapboost.git
cd snapboost
pip install -r requirements.txt
pip install -e .
Releases are published to PyPI via GitHub Actions when a GitHub release is created.
References & Citation
If you use this package or the HNBM framework in research, please cite the original SnapBoost paper:
Thomas Parnell, Andreea Anghel, Małgorzata Łazuka, Nikolas Ioannou, Sebastian Kurella, Peshal Agarwal, Nikolaos Papandreou, and Haralampos Pozidis. SnapBoost: A Heterogeneous Boosting Machine. Advances in Neural Information Processing Systems, 33, 2020.
@inproceedings{parnell2020snapboost,
title = {{SnapBoost}: A Heterogeneous Boosting Machine},
author = {Parnell, Thomas and Anghel, Andreea and {\L}azuka, Ma{\l}gorzata and Ioannou, Nikolas and Kurella, Sebastian and Agarwal, Peshal and Papandreou, Nikolaos and Pozidis, Haralampos},
booktitle = {Advances in Neural Information Processing Systems},
volume = {33},
pages = {20872--20883},
year = {2020},
eprint = {2006.09745},
doi = {10.48550/arXiv.2006.09745}
}
Links: arXiv:2006.09745 · NeurIPS proceedings · IBM Research
For the full bibliography, related heterogeneous-boosting literature (KTBoost, DeepBoost, etc.), and notes on how this repo relates to the original IBM Snap ML implementation, see REFERENCES.md. Additional BibTeX entries are in CITATION.bib.
License
MIT — See LICENSE for full text.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file snapboost-0.1.5.tar.gz.
File metadata
- Download URL: snapboost-0.1.5.tar.gz
- Upload date:
- Size: 16.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d907dbd87b8d6d3ac89535290b7daae9a58bacf0f7d857755c80ab3d1588aea9
|
|
| MD5 |
286a82dd9a95db21155bc7ae636d0f15
|
|
| BLAKE2b-256 |
b42fe21389a0997aebabb9b6ed82c4fc12a3c8ce90f97cf1951dd50f5eb0ea96
|
Provenance
The following attestation bundles were made for snapboost-0.1.5.tar.gz:
Publisher:
python-publish.yml on QianCapital/snapboost
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
snapboost-0.1.5.tar.gz -
Subject digest:
d907dbd87b8d6d3ac89535290b7daae9a58bacf0f7d857755c80ab3d1588aea9 - Sigstore transparency entry: 2333081739
- Sigstore integration time:
-
Permalink:
QianCapital/snapboost@eca1c82d9c03cab5f1c6f728f6900e8c3dbbb1a0 -
Branch / Tag:
refs/tags/v0.1.5 - Owner: https://github.com/QianCapital
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@eca1c82d9c03cab5f1c6f728f6900e8c3dbbb1a0 -
Trigger Event:
release
-
Statement type:
File details
Details for the file snapboost-0.1.5-py3-none-any.whl.
File metadata
- Download URL: snapboost-0.1.5-py3-none-any.whl
- Upload date:
- Size: 8.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
286990993aa2b199edececc071c3cd57a6e205969d79757d331fcb382ea72aa3
|
|
| MD5 |
fa0de3b4c7cab86b2d27f47be4399b5e
|
|
| BLAKE2b-256 |
7dde7d2827790e9d7b09b6479f3879deff559eabc2fbca753b8d2040cc6e9cb3
|
Provenance
The following attestation bundles were made for snapboost-0.1.5-py3-none-any.whl:
Publisher:
python-publish.yml on QianCapital/snapboost
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
snapboost-0.1.5-py3-none-any.whl -
Subject digest:
286990993aa2b199edececc071c3cd57a6e205969d79757d331fcb382ea72aa3 - Sigstore transparency entry: 2333081782
- Sigstore integration time:
-
Permalink:
QianCapital/snapboost@eca1c82d9c03cab5f1c6f728f6900e8c3dbbb1a0 -
Branch / Tag:
refs/tags/v0.1.5 - Owner: https://github.com/QianCapital
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@eca1c82d9c03cab5f1c6f728f6900e8c3dbbb1a0 -
Trigger Event:
release
-
Statement type: