Skip to main content

PrismBoost

PyPI Python License: MIT

PrismBoost is a gradient-boosting classifier/regressor that uses SEFR oblique splits at internal nodes (hyperplane splits instead of axis-aligned thresholds), with an optional fast C++ backend.

from prismboost import PrismBoostClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=0, stratify=y
)

clf = PrismBoostClassifier(random_state=0)   # capacity parameters adapt to the data
clf.fit(X_train, y_train)
print(clf.score(X_test, y_test))
print(clf.auto_config_)                      # what was chosen for this training set

Adaptive defaults

Every capacity parameter defaults to "auto" and is resolved from the training-set shape at fit time, the way CatBoost adapts its learning rate to dataset size. The rules were calibrated on the per-dataset Optuna optima of the 121 PMLB benchmark datasets, so an untuned model starts near a sensible configuration instead of one fixed point that under-fits large data and over-fits small data.

Parameter "auto" rule
n_estimators 200
learning_rate 0.05 below 500 rows, else 0.1 (keeps learning_rate * n_estimators near the tuned optimum)
max_depth 4 below 500 rows, else 6
min_samples_leaf sqrt(n_samples) / 2, clipped to [5, 30]
min_samples_split 2 * min_samples_leaf
subsample 0.8, or 1.0 below 100 rows
split_mode hybrid up to 50 features, else hybrid_sampled

Only the shape of X is used, never y, so folds of equal size resolve identically and no label information leaks. Passing an explicit value disables adaptation for that parameter alone:

clf = PrismBoostClassifier(max_depth=3, random_state=0)  # depth fixed, the rest still auto
clf.fit(X_train, y_train)
clf.max_depth_, clf.n_estimators_   # (3, 200)

Class imbalance is deliberately left alone (class_weight=None), matching XGBoost and CatBoost defaults. To reproduce pre-0.2 behaviour, pass the old values explicitly: n_estimators=100, learning_rate=0.1, max_depth=3, min_samples_leaf=10, min_samples_split=2, subsample=1.0, split_mode="hybrid_sampled".

Leaf regularization (reg_lambda)

reg_lambda is the L2 penalty on leaf weights, the same knob as XGBoost's reg_lambda. It enters the Newton step as sum(w r) / (sum(w h) + reg_lambda) and the split gain as G^2 / (H + reg_lambda).

It defaults to 0.0, so results from earlier versions are unchanged unless you set it. Raising it matters most on imbalanced classification: a nearly pure leaf has h = p (1 - p) close to zero, so the unregularized Newton step is large and the accumulated scores can saturate the softmax. On a 3-class problem at a 2/10/88 class split, 200 rounds:

reg_lambda test log loss
0.0 0.399
1.0 0.233
5.0 0.207
20.0 0.188

(Predicting the class prior scores 0.437 on the same split.) Accuracy-style metrics are much less sensitive to this than log loss is, which is why the PMLB study above, scored on macro-F1 and ROC-AUC, did not surface it.

Measured directly: over 33 PMLB datasets with Optuna tuning, reg_lambda is neutral on ROC-AUC (13 wins, 16 losses, median difference 0.0000). It acts on calibrated probabilities, so tune it when you score with a proper scoring rule such as log loss, or when classes are imbalanced, and expect little from it on ranking metrics.

Early stopping (eval_set, early_stopping_rounds)

Pass validation data to fit and set early_stopping_rounds to stop once the validation loss (log loss for classifiers, squared error for the regressor) has not improved for that many stages. The ensemble is truncated to the best stage, so n_estimators becomes an upper bound:

clf = PrismBoostClassifier(n_estimators=1600, early_stopping_rounds=50)
clf.fit(X_train, y_train, eval_set=(X_val, y_val))
clf.best_iteration_    # stages kept
clf.validation_loss_   # validation loss after each fitted stage

Refitting with n_estimators=clf.best_iteration_ and the same random_state reproduces the early-stopped model exactly. Passing eval_set without early_stopping_rounds only records validation_loss_. On a 4,000-row binary problem with a 1,600-stage cap, early stopping kept 134 stages, cut fit time from 27 s to 3 s, and lowered test log loss from 0.418 to 0.328.

Why oblique boosting?

Axis-aligned GBDTs approximate curved boundaries with staircases. PrismBoost fits linear (oblique) splits, so decision surfaces on non-linear problems are typically smoother.

PrismBoost vs XGBoost probability surfaces on moons

Predicted-probability surfaces on moons: PrismBoost (left) vs XGBoost (right).

Decision boundaries on six synthetic 2D datasets

Decision boundaries on six synthetic 2D datasets (rows) across classifiers (columns). Lower surface roughness S is smoother.

Mean surface roughness by classifier

Benchmark highlights (PMLB)

Evaluated on 121 Penn Machine Learning Benchmark classification datasets against strong baselines (CatBoost, LightGBM, LightGBM-linear, XGBoost, SPORF, Random Forest, Logistic Regression). Hyperparameters are tuned with Optuna; scores are repeated stratified CV.

Median scores (higher is better for F1 / ROC-AUC; lower is better for inference latency):

Model Median macro-F1 Median ROC-AUC Median inference (ms/row)
PrismBoost 0.903 0.976 0.056
CatBoost 0.892 0.976 0.184
XGBoost 0.875 0.970 0.615
LightGBM-linear 0.875 0.970 0.567
LightGBM 0.870 0.972 0.550
Random Forest 0.862 0.964 5.140
SPORF 0.856 0.970 10.058
Logistic Regression 0.822 0.942 0.053

Average ranks (1 = best; Friedman tests significant for macro-F1 and ROC-AUC):

Model Macro-F1 rank ROC-AUC rank
CatBoost 3.33 3.48
PrismBoost 3.88 4.26
LightGBM-linear 3.99 4.26
LightGBM 4.10 4.10
XGBoost 4.31 4.24
Logistic Regression 5.31 5.66
Random Forest 5.48 5.31
SPORF 5.61 4.70

Nemenyi critical-difference diagrams

Nemenyi critical-difference diagrams (α = 0.05). Models connected by a bar are not significantly different.

On these data, PrismBoost is competitive with modern GBDTs on accuracy while remaining among the fastest at inference (second only to logistic regression; fastest non-linear model by median latency).

Install

pip install prismboost

Requires Python 3.10–3.13. A C++17 compiler and CMake are used when building the optional native extension (included for common platforms via wheels / sdist build).

From source (editable / development):

pip install -e ".[dev]"

If the C++ extension fails to build, the package still works via the pure-Python backend.

Extras:

pip install "prismboost[examples]"   # Optuna for the tuning example
pip install "prismboost[dev]"        # pytest, ruff

Public API

Name Description
PrismBoostClassifier Classifier (sklearn-compatible)
PrismBoostRegressor Regressor
SEFR Linear weak learner used inside oblique splits
auto_boosting_config(n_samples, n_features) The "auto" default rules, callable for inspection

Legacy names

The project was previously called SEFRBoost. Those names are aliases of the classes above — the same objects, so isinstance checks and old pickles keep working — and are kept for backwards compatibility:

Legacy name Now
SEFRBoostClassifier / SEFRBoostRegressor PrismBoostClassifier / PrismBoostRegressor
SEFRGradientBoostingClassifier / SEFRGradientBoostingRegressor PrismBoostClassifier / PrismBoostRegressor
prismboost.sefr_gbdt, prismboost.sefr_boost prismboost.prism_boost

SEFR itself is not legacy: it is the linear model that produces each oblique split, and it keeps its name.

Features

  • Oblique tree splits from SEFR (closed-form linear separator)
  • Binary and multiclass classification; regression
  • Newton (second-order) split gain and leaf values, as in XGBoost and CatBoost; second_order=False selects the first-order variant, where the per-sample Hessian is replaced by 1 so the gain becomes variance reduction and leaves hold the mean residual
  • Optional C++ backend for faster fit/predict
  • sklearn estimator API (fit, predict, predict_proba, pipelines, pickling)
  • Works with Optuna / GridSearchCV / RandomizedSearchCV

Examples

pip install "prismboost[examples]"
python examples/quickstart.py
python examples/optuna_tuning.py --n-trials 20

Docs mirror: docs/optuna_tuning.rst.

Tests

pip install -e ".[dev]"
pytest -q

License

This project is licensed under the MIT License.

Third-party note: prismboost._utils includes code derived from wnb under the BSD 3-Clause License.

Citation

If you use PrismBoost in academic work, please cite the accompanying paper (to be updated on publication).

Authors

  • Hamidreza Keshavarz
  • Reza Rawassizadeh

Metadata

Release files for prismboost 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for prismboost 0.4.0
File Size Uploaded
prismboost-0.4.0.tar.gz 754.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for prismboost 0.4.0
File Interpreter ABI Platform
prismboost-0.4.0-cp312-cp312-macosx_26_0_arm64.whl CPython 3.12 CPython 3.12 macOS 26.0+ ARM64 Details

Total release size: 924.3 kB

Release files / prismboost-0.4.0.tar.gz

Download URL prismboost-0.4.0.tar.gz
Size 754.8 kB
Tags Source
SHA-256 checksum
How to use checksums
ae34968c1de6210fe88e105532a5ab2ce07358a96db195fdee347913e80e0e36
BLAKE2b-256 checksum
How to use checksums
9e15014ea47dd4bc77c27d7b5646be159a44e644f6f2dadd738b9af1f972a3c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.0

Release files / prismboost-0.4.0-cp312-cp312-macosx_26_0_arm64.whl

Download URL prismboost-0.4.0-cp312-cp312-macosx_26_0_arm64.whl
Size 169.4 kB
Tags CPython 3.12 macOS 26.0+ ARM64
SHA-256 checksum
How to use checksums
bf3397d717140b4233bd2f24652a63246ba4dbe02ad846e0bb858b181ab18826
BLAKE2b-256 checksum
How to use checksums
5e9bc28a13bce8a7509fed839bbf2ee33e8ba136c6565593dc7c604dfd708a63
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.0

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page