RuleTreeRank
RuleTreeRank (RTR) is an interpretable Learning-to-Rank framework that models ranking as a two-stage process: it first groups items with a shallow rule tree, then refines the score locally through instance-based comparisons.
RTR is designed for query-supported ranking problems. In candidate screening, for example, a query is a job offer and each item is a candidate. The same candidate can be relevant for one job and irrelevant for another, so the final ranking is induced within each query by sorting the RTR scores in descending order.
What RTR Does
- Stage I: rule-based grouping. A shallow
RuleTreeRegressorpartitions the feature space into interpretable leaves and assigns each item a coarse score. - Stage II: local refinement. Inside each leaf, RTR learns an interpretable pairwise distance model and uses a local k-NN correction over the Stage-I residuals.
- Explanations. The first stage gives a decision path, while the second stage exposes the local neighbours and pairwise rules used for the correction.
The package also exposes MixedRTR, a query-aware variant that trains the local refinement by both tree leaf and query when query-specific historical context is required.
Installation
pip install ruletreerank
Latest development version:
pip install git+https://github.com/jacons/RuleTreeRank.git
From a clone, for development:
git clone https://github.com/jacons/RuleTreeRank.git
cd RuleTreeRank
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest
Requires Python 3.10 or newer.
Minimal Usage
import numpy as np
from RuleTree import RuleTreeRegressor
from ltr_utility import ModelParam
from ruletreerank import KNNRegFast, PairwiseDistanceTree, RuleTreeRank
X_train = np.random.default_rng(0).normal(size=(100, 5))
y_train = X_train @ np.array([1.4, -0.7, 0.5, 0.0, 0.2])
q_train = np.repeat(np.arange(10), 10)
model = RuleTreeRank(
distance_f=ModelParam(PairwiseDistanceTree, {
"base_regressor": ModelParam(RuleTreeRegressor, {"max_depth": 2, "random_state": 0}),
"feature_concat": False,
"feature_diff": True,
"feature_sq_diff": True,
"subsample": 0.5,
}),
aggregation_f=ModelParam(KNNRegFast, {"n_neighbors": 5, "n_jobs": 1}),
base_regressor=RuleTreeRegressor(max_depth=2, random_state=0),
dist_objective="residuals",
)
model.fit(X_train, y_train, q_train)
scores = model.predict(X_train, q=q_train, output="full")
ranking_for_query_0 = np.argsort(-scores[q_train == 0])
Mathematical View
Ranking data is represented as a collection of queries and their items:
$$ \mathcal{D} = {(q, X_q, y_q)}_{q \in \mathcal{Q}}, \quad X_q = {\mathbf{x}1,\ldots,\mathbf{x}{n_q}}. $$
RTR learns a two-stage pointwise scoring function:
$$ f(\mathbf{x}) = r(\mathbf{x}) + s(\mathbf{x}). $$
The first term, $r(\mathbf{x})$, is the score assigned by a shallow decision tree. If $\ell(\mathbf{x})$ is the leaf reached by item $\mathbf{x}$, the baseline score is the leaf mean:
$$ r(\mathbf{x}) = \frac{1}{|\ell(\mathbf{x})|} \sum_{\mathbf{x}_j \in \ell(\mathbf{x})} y_j. $$
Training residuals are then computed as $\varepsilon_j = y_j - r(\mathbf{x}_j)$. The second term, $s(\mathbf{x})$, is a local correction: for a test item $\mathbf{x}^\star$, RTR retrieves the $k$ nearest training items in the same leaf using a learned pairwise distance model, then averages their residuals:
$$ s(\mathbf{x}^\star) = \frac{1}{k} \sum_{\mathbf{x}_j \in N_k(\mathbf{x}^\star)} \varepsilon_j. $$
The final ranking for a query is obtained by sorting items by $f(\mathbf{x})$ in decreasing order.
Examples
- examples/min_example.ipynb: minimal RTR example on a scikit-learn dataset, with tables and plots showing the double-stage behaviour.
- examples/synthetic_ranking_task.ipynb: synthetic ranking task with a hidden linear scoring function and relevance labels induced by ranking.
Repository Layout
src/ruletreerank/: core RTR, MixedRTR, pairwise distance tree, and fast k-NN components.src/ltr_utility/: shared interfaces, dataset utilities, query splitting, clustering, and explanation helpers.examples/: executable notebooks for minimal and synthetic ranking workflows.tests/: smoke test covering the two-stage fit/predict path.
Metadata
Release files for ruletreerank 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ruletreerank-0.1.0.tar.gz | 48.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ruletreerank-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 102.2 kB
Release files / ruletreerank-0.1.0.tar.gz
| Download URL | ruletreerank-0.1.0.tar.gz |
|---|---|
| Size | 48.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
30e6fc115cbc33aae4b3322f4b880146b705c47866b66e3bcb9215947f68ae7c
|
|
BLAKE2b-256 checksum How to use checksums |
b8b08bd31cb1f8e835ede14be8e0d027ad155c95d217638c3897c547ba7e68d1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.
Transparency logRelease files / ruletreerank-0.1.0-py3-none-any.whl
| Download URL | ruletreerank-0.1.0-py3-none-any.whl |
|---|---|
| Size | 53.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7157f9fd7d2e85431f8dbfac39a2a39820b2c9b961f8298ecc45bad334ed140f
|
|
BLAKE2b-256 checksum How to use checksums |
a635b775606c361abff6829a74a941355c26160ddb04982cb425810f8c9dbc51
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.
Transparency log