pyterrier-tar
Replay-only technology-assisted-review (TAR) stopping extensions for PyTerrier.
This is an independent package, not a PyTerrier fork. It depends only on
documented public pyterrier APIs so that the dependency can be upgraded and
compatibility tested independently. It does not vendor PyTerrier code.
Status
Alpha (0.1.0). Each stopping rule replays a recorded, labelled ranking and reports where it would have stopped, why, and at what review cost. Rules whose authors released code or results are checked against them, and the rest against their published definitions; PUBLISHED_EVIDENCE.md lists every check and what it does not establish.
The package also registers the public CLEF eHealth TAR collections as
tar:clef2017, tar:clef2018, and tar:clef2019. These dataset adapters
hash-check the public archives and pinned released AutoTAR rankings before
using them for replay validation.
Stopping rules
For resumable experiments, uncertainty and failure reports, ranking comparisons, TARexp/CSV history import, and diagnostic plots, see BENCHMARKING.md.
Every rule is a PyTerrier transformer over a ranked, labelled trajectory
(qid, rank, label, in review order). transform returns the reviewed
prefix; stop_report gives the stop and its cost; stop_trace explains the
decision; calculation_trace shows every checkpoint the rule evaluated. A rule
sees labels only up to the checkpoint it is judging, which the test suite
enforces for all of them.
A rule that aims at a recall target takes target_recall as its first
argument, with no default, and a rule that needs the collection size takes
collection_size next: tar.IPHyperbolic(.8, 3000), tar.CMHHeuristic(.8, 3000).
Control-set rules take their screened frame first: tar.QBCB(control, .8).
Rules and metrics that need the collection size either require
collection_size or accept it optionally. When it is optional (SAFE,
BetaBinomial, stopping_metrics) and omitted, the trajectory is taken to be
the whole collection, so pass it whenever the ranking is truncated.
CLEFTARDataset.complete_ranking appends a topic's unranked documents. A
qid is matched across frames as text, so 1 and '1' are the same query.
Choosing one
tar.catalogue() # every rule: family, promise, what it needs, source
tar.catalogue(promise='certificate') # only the rules that carry a guarantee
Start from what you can supply. With labels alone, you can run any
trajectory rule or the checkpoint rules that also take a collection size. With
calibrated probabilities from your classifier, the estimation rules become
available. If you can screen an independent random sample, QBCB and
TargetRecapture are the only rules here that give a guarantee rather than an
estimate, and they charge that sample as review cost.
What a rule promises
The four groups below say what a rule reads. What it promises is a
different split, and decides how it should be evaluated: heuristics promise
nothing about recall, estimators promise it if their model holds, and
certificates (QBCB, TargetRecapture) promise it with a stated probability
over the draw of their control sample. Use reliability_summary for an
estimator and coverage_summary for a certificate; see
EVALUATION.md.
Trajectory rules — labels only
| Rule | Stops when | Source |
|---|---|---|
Kneedle |
At BMI checkpoints it locates the knee of the gain curve, then compares the pre-knee to post-knee slope ratio against 156 - min(relevant at knee, 150). knee_distance='absolute' (default) matches the released implementations; 'signed' is Satopää et al.'s Kneedle. |
Cormack and Grossman (2016a) |
Budget |
Three quarters of the collection is reviewed, or n >= 10N/r with a knee slope ratio of at least 6. |
Cormack and Grossman (2016a) |
Rule2399 |
Reviewed documents reach 2399 + 1.2 x relevant found. |
Cormack and Grossman (2016b) |
ReviewHalf |
Half the known collection is reviewed. | Yang and Lewis (2022), TARexp |
FixedRound |
A set number of review batches is done. | Yang and Lewis (2022), TARexp |
BatchPrecision |
patience consecutive batches have precision at or below the cutoff. |
Yang, Lewis, and Frieder (2021) |
ConsecutiveIrrelevant |
A run of count non-relevant documents completes. |
common heuristic |
SAFE |
Every supplied key paper is found, at least twice the seed positives and min_fraction of the collection are reviewed, and the last consecutive documents are non-relevant. |
Boetje and van de Schoot (2024) |
Oracle |
The prefix first reaches the target recall, using complete labels. An offline lower bound, never deployable. | — |
Checkpoint rules — a statistical test at fixed intervals
| Rule | Stops when | Source |
|---|---|---|
CMHHeuristic |
The biased-urn (BUSCAR) test rejects "recall is still below target" at alpha. |
Callaghan and Müller-Hansen (2020) |
AnytimeCMH |
The same test with an alpha / (k(k+1)) spending schedule across checkpoints, so the levels sum to alpha. Conditional on the urn p-values being calibrated for the recorded design. |
original to this package (A. Fletcher) |
PoissonPoint (IP-P) |
A power-law rate fitted to window relevance yields a Poisson upper bound on the documents still unfound, and found >= target x (found + bound); also when the later windows contain nothing relevant. |
Stevenson and Bin-Hezam (2023) |
IPHyperbolic (IP-H) |
As IP-P with a hyperbolic rate. tail='released' (default) reproduces the released code's expected-remaining formula, which understates it by (1-b)^2; tail='model' integrates the fitted rate. |
Stevenson and Bin-Hezam (2023) |
Control-set rules — an independently screened sample
| Rule | Stops when | Source |
|---|---|---|
QBCB |
The Jth positive control appears in the ranking, where J is the smallest rank whose binomial bound certifies the target. |
Lewis, Yang, and Frieder (2021) |
TargetRecapture |
All k = ceil(-ln(1-confidence)/(1-target)) sampled targets have been reviewed. |
Cormack and Grossman (2016a) |
BaselineInclusionRate |
Relevant documents found reach target x the prevalence estimated from a random pilot. It never fires when the pilot finds nothing. |
pilot-prevalence heuristic |
sample_control |
Helper: draws an unlabelled worklist to screen independently. Control screening is charged in stop_report and stopping_metrics. |
— |
Estimation rules — recorded probabilities or sampling metadata
| Rule | Stops when | Source |
|---|---|---|
Quant, QuantCI |
Estimated recall from calibrated probabilities reaches the target; QuantCI first subtracts nstd standard deviations. Pass score_snapshots to replay per-round model scores instead of one fixed column. |
Yang, Lewis, and Frieder (2021) |
AutoStop |
A Horvitz-Thompson total over a fixed with-replacement sample certifies the target, under loose, strict_v1, or strict_v2; strict_v2 needs collection_size. Not the adaptive procedure. |
Li and Kanoulas (2020) |
SCAL |
The prefix's Horvitz-Thompson estimate reaches the target share of the estimate over the whole recorded sample. The sample past the cutoff is charged as review cost; pass control_documents() to stopping_metrics so it counts towards recall too. |
Cormack and Grossman (2016b) |
Chao |
The Chao1 estimate of unseen relevant documents satisfies found >= target x (found + unseen). Simpler than the paper's multi-model estimators. |
Bron et al. (2025) |
BetaBinomial |
The beta-binomial posterior that few enough relevant documents remain exceeds confidence. Assumes the unreviewed tail is exchangeable with the reviewed prefix, so it errs late on a priority ranking. |
original to this package (A. Fletcher) |
EVPI, EVPIGreedy, EVPISmooth, EVPIBatch |
The expected value of perfect information about the next document (or window) falls to the screening cost. The document that triggers the stop is included in the reviewed prefix. | original to this package (A. Fletcher) |
AnytimeCMH, BetaBinomial, and the EVPI family are original to this
package rather than replays of published methods, so they carry no parity
claim: they are covered by formula tests and the shared leakage and trace
checks only.
Learned rule
GRLStop trains a PPO policy over 100 review windows, where the state is the
relevance rate in each reviewed window, a logistic-regression prediction for
each unreviewed one, the current window, and the target recall. It needs the
grl extra. Source: Bin-Hezam and Stevenson (2025).
Datasets and helpers
CLEFTARDataset backs pt.get_dataset('tar:clef2017') and its 2018/2019
siblings: hash-pinned public archives, one candidate pool per topic,
complete_ranking() to append unretrieved documents, and label_ranking() to
attach qrels only after a ranking exists.
Metrics
stopping_metrics reports recall, cost, review fraction, reliability, oracle
depth, and the CLEF loss_e/loss_r/loss_er components, charging any
separately screened controls once. stopping_frontier summarises several
targets, and reliability and review_fraction are ir-measures metrics for
pt.Experiment.
reliability_summary gives the share of topics reaching the target with a
Wilson interval, which is the check for an estimator. coverage_summary gives
the share of repeated control draws that certified it, which is the check for a
certificate.
Adding a rule
One file per approach, inside the folder for what the rule reads:
rules/trajectory/ for labels only, rules/checkpoint/ for a test at fixed
checkpoints, rules/control/ for an independently screened sample, and
rules/estimation/ for recorded probabilities or sampling metadata.
# src/pyterrier_tar/rules/trajectory/my_rule.py
from .._base import _TrajectoryStoppingRule, _batch_positions
class MyRule(_TrajectoryStoppingRule):
"""Stops once a batch holds no relevant documents."""
_equation = 'fire if the batch ending at n has no relevant document'
def __init__(self, batch_size: int = 200, initial_documents: int = 1):
if batch_size < 1 or initial_documents < 1:
raise ValueError('batch sizes must be positive')
self.batch_size = batch_size
self.initial_documents = initial_documents
def _checkpoint_rows(self, ranked, labels):
for stop in _batch_positions(len(labels), self.batch_size, self.initial_documents):
batch = labels[max(0, stop - self.batch_size):stop]
yield {'index': stop, 'relevant_in_batch': int((batch > 0).sum()),
'fired': not (batch > 0).any()}
_checkpoint_rows yields one dict per checkpoint, with at least index and
fired; the base class turns it into the stop, stop_report, stop_trace, and
calculation_trace, so those cannot disagree. A rule that is not
checkpoint-shaped implements _stop_position(labels) instead, and one that
needs other columns implements _stop_position_for_results(ranked, labels).
Add _trace_details for the columns stop_trace should carry.
The row for checkpoint k must use only labels[:k]. Export the class from
the group's __init__.py, from rules/__init__.py, and from
pyterrier_tar/__init__.py; the shared tests then pick it up, including the
leakage check and the report/trace consistency check. New rules are expected to
come with a test that re-derives the stop from the rule's own definition, as
tests/test_rule_validity.py does for the others.
Scope and safety boundary
pyterrier-tar replays recorded, labelled trajectories. It is not a live active-learning or screening controller, and it does not certify target recall in deployment. Labels beyond the reviewed prefix must not be exposed to a stopping rule. Separately sampled controls must be independently screened and accounted for in review cost.
Install
python -m pip install pyterrier-tar
python -m pip install 'pyterrier-tar[grl]'
The grl extra installs gymnasium, scikit-learn, and stable-baselines3
for GRLStop; tutorial adds PyTerrier's Java support and matplotlib for the
notebooks that build Terrier indexes.
Quick start
import numpy as np
import pandas as pd
from pyterrier import tar
# A recorded review: ranked documents, with the label revealed as each is read.
# Relevance thins out down the ranking, as it does in a real screening run.
rng = np.random.default_rng(0)
size = 6000
prevalence = .7 * np.exp(-np.arange(size) / 300) + .002
trajectory = pd.DataFrame({
'qid': 'CD008081',
'docno': [f'd{position}' for position in range(size)],
'rank': range(size),
'label': (rng.random(size) < prevalence).astype(int), # 1 for relevant
})
rule = tar.Kneedle()
reviewed = rule.transform(trajectory) # the prefix the rule would have read
report = rule.stop_report(trajectory) # stop, fired, review cost
print(rule.stop_trace(trajectory).iloc[0].reason)
print(rule.calculation_trace(trajectory, index=int(report.stop.iloc[0])).tail())
qrels = trajectory.loc[trajectory.label > 0, ['qid', 'docno', 'label']]
print(tar.stopping_metrics(trajectory, reviewed, qrels, target_recall=.95, report=report))
On this trajectory Kneedle stops after 2,001 of 6,000 documents, having found
94.7% of the relevant ones; stop_trace says the knee slope ratio reached its
dynamic threshold, and calculation_trace shows the ratio at every checkpoint
it tested.
Development
python -m pip install -e '.[test,grl,tutorial]'
pytest
test runs the suite, grl adds the GRLStop tests, and tutorial the
notebooks. On a CUDA-free machine, UV_TORCH_BACKEND=cpu keeps grl from
pulling the GPU wheels.
Results on CLEF
RESULTS.md shows how each rule performs on the CLEF 2017, 2018,
and 2019 TAR collections: reliability and review cost at recall targets 0.8,
0.9, and 0.95, per collection, with rules grouped by what they promise rather
than ranked. It describes one ranker, the released AutoTAR runs; a rule that
reads the ranking can do much worse behind a weaker one.
RESULTS_UNCERTAINTY.md gives the intervals behind each
row and the largest individual shortfalls. examples/results_table.py
regenerates both.
Published Evidence
The exact public checks and remaining non-parity boundaries are listed in PUBLISHED_EVIDENCE.md. The examples under examples/ are intentionally separate from CI when they require network access, large public archives, or locally held trajectories.
For real-data or external-source smoke checks:
python examples/published_cmh_buscar.py
python examples/published_clef_tar_eval_metrics.py
python examples/published_ip_h_clef.py
python examples/published_kneedle_clef2017.py
python examples/published_point_process_clef2017.py
python examples/published_autostop_clef2017.py
PYTHONPATH=examples python examples/published_baselines_clef.py
Run them from a scratch directory: each caches its inputs in data/ relative
to the working directory, about 1 GB in total.
examples/notebooks/parity_published.ipynb runs all of them and shows one
parity table. Every notebook is either a demo_ (how the package behaves) or a
parity_ (a published result reproduced); see examples/notebooks.md.
published_kneedle_clef2017.py checks Kneedle's released 86,243 effort. The
same seven scripts run weekly in CI's published-parity job. The notebooks that
build Terrier indexes need the tutorial extra, which installs PyTerrier's
Java support.
Only the metrics asserted by each evidence check should be treated as validated.
For example, the CLEF metric checker asserts only the official recall, total
cost, and loss fields exposed by stopping_metrics(), while the IP-H CLEF
checker asserts recall, cost, reliability, loss, and relative-error parity
across all nine CLEF 2017--19 settings.
Known gaps
GAPS.md lists what this package does not establish: the replay-versus-live question, published results whose inputs were never released, methods named but not implemented, and release chores. New findings about limits belong there rather than in a commit message.
Compatibility policy
The supported PyTerrier range is declared in pyproject.toml. New code must import only public pyterrier names; imports from private modules such as pyterrier._ops are prohibited. Each supported PyTerrier release will receive a focused compatibility test before its range is widened.
License and notices
This source is licensed under MPL-2.0; see LICENSE.txt. PyTerrier-derived source and the IP-H implementation boundary are recorded in THIRD_PARTY_NOTICES.md.
Metadata
Release files for pyterrier-tar 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pyterrier_tar-0.1.0.tar.gz | 202.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pyterrier_tar-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 303.6 kB
Release files / pyterrier_tar-0.1.0.tar.gz
| Download URL | pyterrier_tar-0.1.0.tar.gz |
|---|---|
| Size | 202.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
de9ebd883bf63e8e4f3fe211d80ef50e721268372aa0f2317ccc9adb67396a1b
|
|
BLAKE2b-256 checksum How to use checksums |
3c86ead3f0b64bd8e809aa29fe68d2511b0e06026b0a451b992fedf9f398464b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / pyterrier_tar-0.1.0-py3-none-any.whl
| Download URL | pyterrier_tar-0.1.0-py3-none-any.whl |
|---|---|
| Size | 101.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2ca6ec956d8f48f7bf0ce0ed2c5fbc5655e5f9c18defe02014ced711268e54d9
|
|
BLAKE2b-256 checksum How to use checksums |
63f3e74b7773a7a179277f73dcfb3cb7214442913777df6e0a9e2791345c9d67
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log