Skip to main content

TreeIG

PyPI version

TreeIG computes exact Integrated Gradients for tree-based models. It decomposes the change in a fitted tree model's scalar output between a baseline input $x_0$ and an observation $x$ into additive feature contributions.

For each observation, TreeIG returns feature attributions $\phi_j$ satisfying

$$\sum_j \phi_j = F(x) - F(x_0),$$

where $F$ is the scalar model output being explained. For regression models, $F$ is the prediction. Exact classification backends use native raw margins. TreeIGNumeric can additionally transform probability-only classifiers to binary log odds or centered multiclass log scores.

Integrated Gradients (Sundararajan, Taly, and Yan, 2017) defines feature attributions by integrating model gradients along a straight-line path from a baseline $x_0$ to the observation $x$.

At first glance, Integrated Gradients appears mismatched with piecewise-constant tree models: gradients vanish almost everywhere and are undefined at split boundaries. Hentschel (2026) shows that, for tree-based models, the path-integral of the gradients reduces to the sum of prediction jumps at split boundaries crossed along the integration path. The resulting attribution is exact — no Monte Carlo sampling, no numerical quadrature, no approximation parameters.

Because TreeIG replaces numerical quadrature and sampling with a finite sum over split crossings, it is fast in practice. For many real-world models — hundreds of trees, hundreds of features — attribution over thousands of observations completes in a few milliseconds on a modern laptop. (See the example notebook for timings.) For many typical use cases TreeIG is competitive with, and often faster than, TreeSHAP, which is itself considered fast.

TreeIG also includes TreeIGNumeric, a model-agnostic fallback that recovers the same crossing-sum attribution through numerical event detection when exact structural support is unavailable.

Recommended baseline construction

For Integrated Gradients, the baseline determines the prediction contrast being explained. CBaseline is the preferred way to construct TreeIG baselines. It produces empirical, prediction-neutral baseline distributions whose weighted mean model output is the chosen reference prediction. TreeIG then explains the model prediction relative to that reference level rather than relative to an arbitrary feature vector such as the feature-wise mean.

TreeIG accepts a CBaseline Background directly and evaluates its weighted baseline paths efficiently. See CBaseline for construction choices and the interpretation of the reference prediction f0.

Installation

pip install treeig
pip install cbaseline  # recommended baseline construction
pip install "treeig[catboost]"  # optional numerical CatBoost support

Requires Python ≥ 3.9, NumPy, and Numba. Model backends (scikit-learn, XGBoost, LightGBM, and CatBoost) are not installed automatically; install whichever you use.

Quickstart

import numpy as np
import treeig as tig

# model is a fitted supported tree model
x0 = X_train.mean(axis=0)
X_eval = X_test[:100]

ig = tig.TreeIG(model, baseline=x0)
phi = ig.attribute(X_eval)

For libraries integrating TreeIG, the public adapter surface also includes:

tig.supports(model)                       # exact backend availability
ig.model_output(X_eval)                   # scalar output being attributed

numeric = tig.TreeIGNumeric(model, baseline=x0)
numeric.model_output(X_eval)              # numeric backend output scale

The single-vector example above is the minimal API. For substantive attribution, prefer a prediction-neutral distribution constructed with CBaseline.

Weighted baseline distributions are first-class baselines. Pass either a matrix and aligned weights or a CBaseline Background directly:

ig = tig.TreeIG(model, baseline=background)  # uses .rows and .weights
phi = ig.attribute(X_eval)

# Equivalent explicit form; weights are normalized internally.
phi = ig.attribute(
    X_eval,
    baseline=background.rows,
    baseline_weights=background.weights,
    baseline_batch_size=25,
)

TreeIG preserves each baseline-specific path and performs the weighted aggregation inside a compiled loop. With return_by_baseline=True, attribute returns (weighted, by_baseline) for diagnostics.

Compiled baseline traversal is also used by loss_attribution and multiclass_loss_attribution. For multiclass log loss, TreeIG merges class-score events in chronological path order before applying each softmax-loss change. Pass the complete baseline distribution in one call instead of invoking the explainer once per baseline: this amortizes tree parsing and model dispatch and keeps baseline aggregation inside compiled code.

phi has the same shape as X_eval. Row i, column j is the contribution of feature j to the model-output change from x_0 to X_eval[i].

For regression models, the completeness property holds exactly:

np.testing.assert_allclose(
    phi.sum(axis=1),
    model.predict(X_eval) - model.predict(x0.reshape(1, -1))[0],
)

Why TreeIG?

Standard Integrated Gradients defines feature contributions by integrating model gradients along a straight-line path from a baseline input to the observation. Tree models are piecewise constant, so ordinary gradients are zero almost everywhere and undefined at split boundaries.

TreeIG uses the tree structure directly. Along the interpolation path

$$ x(t) = x_0 + t,(x - x_0),\qquad 0 \le t \le 1, $$

a tree prediction changes only when the path crosses a split threshold. TreeIG finds those crossings exactly and assigns each prediction jump to the feature responsible for the crossing. For ensembles, contributions are summed across trees. The result is an exact additive decomposition without numerical quadrature.

The distributional-derivative perspective makes this precise. Along the interpolation path the prediction is piecewise constant, and its generalized derivative is a sum of localized impulses at split crossings. The path integral of each impulse is exactly the prediction jump at that crossing.

The top panel shows a step in the tree prediction along the interpolation path. The middle panel shows the corresponding distributional derivative: zero everywhere except at the split crossing. (Here, $\delta(t - t^\ast)$ is the Dirac delta distribution centered at $t^\ast$.) The bottom panel shows that the path integral localizes exactly at the crossing and recovers the prediction jump. TreeIG exploits the fact that integrated gradients applied to trees requires neither numerical differentiation nor numerical integration; it reduces to a simple sum of prediction steps along the integration path $x(t)$.

Standard numerical Integrated Gradients methods try to approximate these impulses using dense interpolation grids. TreeIG instead computes the split-crossing contributions analytically from the fitted tree structure. In this sense, TreeIG plays a role analogous to automatic differentiation for smooth models: rather than numerically searching for discontinuities, it uses the model's computational structure to evaluate the attribution integral exactly and efficiently. (The analogy understates the gain. Automatic differentiation removes derivative approximation but not the numerical quadrature used by Integrated Gradients. TreeIG exploits tree structure to evaluate the attribution integral itself exactly.)

Supported models

TreeIG currently supports tree models with finite numeric feature inputs.

Regression

  • sklearn.tree.DecisionTreeRegressor
  • sklearn.ensemble.RandomForestRegressor
  • sklearn.ensemble.ExtraTreesRegressor
  • sklearn.ensemble.GradientBoostingRegressor
  • xgboost.XGBRegressor
  • xgboost.Booster
  • lightgbm.LGBMRegressor
  • lightgbm.Booster

Classification (raw margins/logits only)

  • sklearn.ensemble.GradientBoostingClassifier
  • xgboost.XGBClassifier
  • lightgbm.LGBMClassifier

For classification models, exact structural TreeIG attributes raw margins or logits. Its current exact engine does not transform ensemble probabilities.

TreeIG computes exact path decompositions directly from the fitted tree structure. Since tree representations differ substantially across implementations, each model family requires customized parsing and routing logic.

Exact support not currently available

The exact TreeIG parser does not currently support:

  • CatBoost;
  • categorical splits;
  • missing-value routing (use feature augmentation for missingness);
  • structurally exact transformed-probability attribution;
  • probability-averaging or vote-share classifiers such as DecisionTreeClassifier, RandomForestClassifier, and ExtraTreesClassifier (because they produce probabilities, not scores).

Many of these can still be attributed with the model-agnostic TreeIGNumeric, described below.

TreeIGNumeric

TreeIGNumeric is a model-agnostic fallback that recovers the crossing-sum attribution by numerically detecting prediction discontinuities along the integration path. It requires no access to model internals — only repeated evaluations of the prediction function — so it applies to many piecewise-constant models the exact parser does not support. Whenever a supported backend is available, exact TreeIG should be preferred.

TreeIGNumeric scans a numerical grid along the integration path to locate changes in the prediction. It then bisects only the changed intervals, four adaptive levels by default, before using local axis-aligned probes to attribute each step to a feature. It preserves completeness for the detected changes and typically produces attributions very similar to exact TreeIG. Because it locates crossings numerically, multiple nearby crossings may occasionally be merged into a single event; exact TreeIG avoids this by enumerating crossings directly from the tree structure.

The defaults grid_size=1024 and max_refine=4 are a practical balance, not an accuracy guarantee. Refinement concentrates additional evaluations around detected changes without increasing the global grid. Completeness can remain exact when offsetting events are hidden inside one coarse interval because the detected jumps still telescope. For allocation-sensitive work, rerun a representative subset at a larger grid and compare the feature attributions themselves. The reproducible probability-forest stress benchmark performs this resolution check against an independent structural crossing oracle.

Two caveats on coverage:

  • CatBoost and other encoded models. TreeIGNumeric removes the parsing barrier, but not the modeling one: interpolating a native categorical feature along the straight-line path is not meaningful, which is a property of Integrated Gradients itself, not of the implementation. TreeIGNumeric works on CatBoost (and similar) models with numeric or one-hot-encoded inputs.
  • Probability-averaging classifiers. By default, TreeIGNumeric retains its original behavior and explains one class probability. With probability_to_score=True, it instead explains binary log odds or a centered multiclass log score. Zero probabilities require an explicit probability_floor; TreeIGNumeric never clips them silently.
import treeig as tig

ig = tig.TreeIGNumeric(model, baseline=x0)
phi, infos, summary = ig.explain(X_eval)
output = ig.model_output(X_eval)

print(summary["mean_abs_residual"])

For a probability-only classifier:

ig = tig.TreeIGNumeric(
    model,
    baseline=x0,
    target=2,                    # omit for binary positive-class log odds
    probability_to_score=True,
    probability_floor=1e-6,     # explicit because tree probabilities may be 0
)
phi = ig.attribute(X_eval)
score = ig.model_output(X_eval)

For binary probabilities, the explained score is log(p1) - log(p0). For K classes, target k selects log(p_k) - mean(log(p)). Pairwise differences are therefore invariant log odds. The floor is applied to every class probability and the probabilities are renormalized before taking logarithms.

These are the canonical scores implied by the complete probability vector: softmax of the centered log scores recovers the original probabilities. They do not reconstruct an unavailable training-time margin. TreeIGNumeric treats the derived score itself as the explicitly defined scalar model output and attributes its jumps after the ensemble probabilities have been aggregated.

Diagnostics

Use explain when you want attributions together with completeness diagnostics.

ig = tig.TreeIG(model, baseline=x0)
phi, infos, summary = ig.explain(X_eval)

print(summary)

Each entry in infos contains diagnostics for one observation:

{
    "n_events":        ...,   # number of split-crossing events
    "endpoint_delta":  ...,   # F(x) - F(x0)
    "attribution_sum": ...,   # sum_j phi_j
    "residual":        ...,   # attribution_sum - endpoint_delta
    "abs_residual":    ...,
}

TreeIGNumeric returns the same fields plus n_coincident_events, the number of events that remained unresolved and were allocated by the fallback rule. It also reports n_refined_intervals, max_refinement_depth, and n_unresolved_intervals. The summary dictionary aggregates these refinement, residual, and event-count diagnostics.

Classification targets

For binary additive-score classifiers, target=None and target=1 both attribute the positive-class margin. target=0 attributes the negative margin, implemented as the negative of the positive-class margin.

ig = tig.TreeIG(model, baseline=x0, target=1)
phi_pos = ig.attribute(X_eval)

ig = tig.TreeIG(model, baseline=x0, target=0)
phi_neg = ig.attribute(X_eval)

For multiclass classifiers, pass the class index explicitly.

ig = tig.TreeIG(model, baseline=x0, target=2)
phi_class_2 = ig.attribute(X_eval)

Exact TreeIG attributes raw class margins. TreeIGNumeric can use the explicit probability-derived score convention above when no native margin exists.

Functional interface

TreeIG also provides a direct functional interface.

phi, infos, summary = tig.compute(
    model,
    baseline=x0,
    X=X_eval,
)

Warmup

TreeIG uses Numba for fast parallel attribution kernels. The first call includes JIT compilation. You can compile in advance with warmup:

ig = tig.TreeIG(model, baseline=x0).warmup(X_eval[:3])
phi = ig.attribute(X_eval)

Subsequent calls on the same model are fast. Attribution for thousands of observations on a typical ensemble completes in well under a second after warmup.

Numerical conventions

TreeIG follows each backend's split-routing convention as closely as possible.

  • scikit-learn trees route left when x[j] <= threshold;
  • LightGBM numeric splits route left when x[j] <= threshold;
  • XGBoost numeric splits route left when x[j] < threshold using float32-style comparisons.

Inputs must be finite numeric arrays. Missing-value routing is not currently implemented, so NaN and Inf values raise errors.

Baselines

The baseline $x_0$ defines the reference point for the decomposition. Common choices include the training-sample mean, a median or representative observation, a domain-specific neutral input, or a fixed benchmark case.

The attribution always explains the difference between the model output at the observation and the model output at the chosen baseline. Different baselines answer different questions.

Interpretation

For an observation $x$, TreeIG reports how much each feature contributes to moving the model output from $F(x_0)$ to $F(x)$ along the straight-line path from $x_0$ to $x$. Positive contributions increase the scalar output relative to the baseline; negative contributions decrease it. The contributions are additive by construction.

Relation to SHAP and TreeSHAP

TreeIG and TreeSHAP answer different attribution questions and generally produce different decompositions. Neither dominates the other.

TreeIG answers: "How much does feature $j$ contribute to the change in prediction as we move continuously from baseline $x_0$ to observation $x$?" The attribution is the integral of partial derivatives along the path from $x_0$ to $x$, which for piecewise-constant trees reduces exactly to a sum of prediction jumps at the split boundaries crossed along the path.

TreeSHAP answers: "How much does feature $j$ shift the expected prediction, averaged over all possible subsets of the other features?" The attribution is an average of discrete inclusion effects, where absent features are marginalized out over a background dataset. There is no path; the reference is the expected prediction over the background distribution.

The two differ in two ways. First, TreeIG takes a specific baseline input $x_0$ as its reference, while TreeSHAP uses a background distribution. Second, TreeIG measures contributions through calculus — integrating how the prediction changes as features move continuously from their baseline values — while TreeSHAP measures them through discrete feature inclusion, asking how much each feature changes the expected prediction when it enters a coalition.

The practical consequence is one of scope. SHAP's coalition construction is indifferent to the prediction surface between the background and the observation: a feature is either in the coalition or out, so the attribution is built from discrete switches and explores a wide neighborhood of hybrid inputs, many far from any natural path between real observations. IG instead follows a single path and accumulates exactly the prediction changes along it, evaluating the model only at convex combinations of two real inputs. SHAP explores a neighborhood; IG traces a path. SHAP's breadth gives sensitivity to model behavior across many feature combinations; IG's specificity gives a precise account of one trajectory through input space.

For a linear model with independent features and $x_0$ equal to the background mean, TreeIG and interventional SHAP coincide. (A linear model is not a tree, so the comparison is to SHAP generally rather than to TreeSHAP.) As the model becomes more nonlinear or the baseline $x_0$ diverges from the background distribution, the two increasingly disagree — reflecting genuine differences in the questions they answer rather than errors in either method.

Examples

XGBoost regression

import numpy as np
import xgboost as xgb
import treeig as tig

model = xgb.XGBRegressor(
    n_estimators=100,
    max_depth=3,
    learning_rate=0.05,
    objective="reg:squarederror",
    random_state=0,
)
model.fit(X_train, y_train)

x0 = X_train.mean(axis=0)
X_eval = X_test[:100]

ig = tig.TreeIG(model, baseline=x0).warmup(X_eval[:3])
phi, infos, summary = ig.explain(X_eval)

print(phi.shape)
print(summary["max_abs_residual"])

Multiclass classification margins

import lightgbm as lgb
import treeig as tig

model = lgb.LGBMClassifier(...)
model.fit(X_train, y_train)

x0 = X_train.mean(axis=0)
X_eval = X_test[:100]

# Attribute class-2 raw margin
ig = tig.TreeIG(model, baseline=x0, target=2)
phi = ig.attribute(X_eval)

Model-agnostic attribution

import treeig as tig

ig = tig.TreeIGNumeric(model, baseline=x0)
phi, infos, summary = ig.explain(X_eval)

print(summary["mean_abs_residual"])

Project status

TreeIG is production-ready for exact attribution of supported tree models in raw-output space. The current release covers the dominant tree ensemble backends in the Python ecosystem. TreeIGNumeric provides a model-agnostic fallback for unsupported piecewise-constant models.

Future extensions may include:

  • exact structural support for CatBoost and other currently unsupported tree implementations;
  • customized handling of categorical split structures and missing-value routing;
  • alternative allocation rules for simultaneous multi-feature effects at coincident crossings.

Citation

If you use TreeIG in your work, please cite:

@misc{hentschel2026treeig,
  author = {Hentschel, Ludger},
  title  = {{TreeIG}: Exact Integrated Gradients for Tree-Based Models},
  year   = {2026},
  url    = {https://www.ludgerhentschel.com/PDFs/Hentschel%20'26g.pdf},
}

License

TreeIG is released under the terms in LICENSE.

References

TreeIG:

Integrated Gradients:

  • Sundararajan, Mukund, Ankur Taly, and Qiqi Yan. 2017. "Axiomatic Attribution for Deep Networks." International Conference on Machine Learning (ICML).

SHAP and TreeSHAP:

  • Lundberg, Scott M., and Su-In Lee. 2017. "A Unified Approach to Interpreting Model Predictions." Advances in Neural Information Processing Systems (NeurIPS).

  • Lundberg, Scott M., Gabriel Erion, and Su-In Lee. 2020. "From Local Explanations to Global Understanding with Explainable AI for Trees." Nature Machine Intelligence.

Popular implementations of Integrated Gradients for smooth models:

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

treeig-0.1.11.tar.gz (58.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

treeig-0.1.11-py3-none-any.whl (45.8 kB view details)

Uploaded Python 3

File details

Details for the file treeig-0.1.11.tar.gz.

File metadata

  • Download URL: treeig-0.1.11.tar.gz
  • Upload date:
  • Size: 58.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for treeig-0.1.11.tar.gz
Algorithm Hash digest
SHA256 b94f9f3b933cf36214e1b31d6092576bb04d7d8650020e16178fd37fe85733b8
MD5 b96b4afefd5d3adba390b8ff6126bad5
BLAKE2b-256 11f965063c875178551831d45594f7a46098d40a807f238137f71af5f6d0e090

See more details on using hashes here.

Provenance

The following attestation bundles were made for treeig-0.1.11.tar.gz:

Publisher: release.yml on LudgerHentschel/treeig

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file treeig-0.1.11-py3-none-any.whl.

File metadata

  • Download URL: treeig-0.1.11-py3-none-any.whl
  • Upload date:
  • Size: 45.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for treeig-0.1.11-py3-none-any.whl
Algorithm Hash digest
SHA256 89f3cec30f954b2ab1d191763eba01b28883e4a2c3e597b473e9cc07ed091fd4
MD5 124e0a1dd0d97de25d21c4b3a81482d4
BLAKE2b-256 8260b8e053ee772e38d928c5f205fb666ddfb8fa4369a9ec438f197d5c6de0dd

See more details on using hashes here.

Provenance

The following attestation bundles were made for treeig-0.1.11-py3-none-any.whl:

Publisher: release.yml on LudgerHentschel/treeig

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.14

2 files

This release

0.1.11 This release

2 files

0.1.10

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page