Skip to main content

pyterrier-tar

Replay-only technology-assisted-review (TAR) stopping extensions for PyTerrier.

This is an independent package, not a PyTerrier fork. It depends only on documented public pyterrier APIs so that the dependency can be upgraded and compatibility tested independently. It does not vendor PyTerrier code.

Status

Alpha (0.1.1). Each stopping rule replays a recorded, labelled ranking and reports where it would have stopped, why, and at what review cost. Rules whose authors released code or results are checked against them, and the rest against their published definitions; the Published Evidence section below summarises every check and what it does not establish.

The package also registers the public CLEF eHealth TAR collections as tar:clef2017, tar:clef2018, and tar:clef2019. These dataset adapters hash-check the public archives and pinned released AutoTAR rankings before using them for replay validation.

File names such as BENCHMARKING.md and paths under examples/ refer to the source distribution, which holds the full documents, the example scripts, and the notebooks. The wheel that pip install uses holds only the package; the Install section shows how to fetch the rest.

Stopping rules

For resumable experiments, uncertainty and failure reports, ranking comparisons, TARexp/CSV history import, and diagnostic plots, see the Benchmarking section below.

Every rule is a PyTerrier transformer over a ranked, labelled trajectory (qid, rank, label, in review order). transform returns the reviewed prefix; stop_report gives the stop and its cost; stop_trace explains the decision; calculation_trace shows every checkpoint the rule evaluated. A rule sees labels only up to the checkpoint it is judging, which the test suite enforces for all of them.

A rule that aims at a recall target takes target_recall as its first argument, with no default, and a rule that needs the collection size takes collection_size next: tar.IPHyperbolic(.8, 3000), tar.CMHHeuristic(.8, 3000). Control-set rules take their screened frame first: tar.QBCB(control, .8).

Rules and metrics that need the collection size either require collection_size or accept it optionally. When it is optional (SAFE, BetaBinomial, stopping_metrics) and omitted, the trajectory is taken to be the whole collection, so pass it whenever the ranking is truncated. CLEFTARDataset.complete_ranking appends a topic's unranked documents. A qid is matched across frames as text, so 1 and '1' are the same query.

Choosing one

tar.catalogue()                        # every rule: family, promise, what it needs, source
tar.catalogue(promise='certificate')   # only the rules that carry a guarantee

Start from what you can supply. With labels alone, you can run any trajectory rule or the checkpoint rules that also take a collection size. With calibrated probabilities from your classifier, the estimation rules become available. If you can screen an independent random sample, QBCB and TargetRecapture are the only rules here that give a guarantee rather than an estimate, and they charge that sample as review cost.

What a rule promises

The four groups below say what a rule reads. What it promises is a different split, and decides how it should be evaluated, because what a rule promises decides what would falsify it. tar.catalogue() gives the promise of each rule.

Promise Rules What is promised The check
Heuristic Kneedle, Budget, Rule2399, ReviewHalf, FixedRound, BatchPrecision, ConsecutiveIrrelevant, SAFE nothing about recall: they detect a flattening gain curve or exhaust a budget the cost/recall trade-off, and excess cost over the oracle depth
Estimator CMHHeuristic, AnytimeCMH, PoissonPoint, IPHyperbolic, BaselineInclusionRate, Quant, QuantCI, AutoStop, SCAL, Chao, BetaBinomial, the EVPI family, GRLStop recall at or above the target if its model holds reliability_summary: the share of topics reaching the target, with a Wilson interval, against the nominal level
Certificate QBCB, TargetRecapture recall at or above the target with probability 1 - alpha over the draw of the control sample coverage_summary: the share of repeated draws that certified the target, for one review

The unit of repetition is the difference that matters. An estimator's promise is about topics, so it is measured across them. A certificate's promise is about the sample it drew, so it is measured by drawing again: coverage across replications of one review. Measuring a certificate across topics answers a question nobody asked.

Costs are comparable only when the control screening is charged. stop_report and stopping_metrics already count it once; a certificate that costs more than a heuristic is buying a guarantee the heuristic never offers. EVALUATION.md is the same guidance as a single page.

Trajectory rules — labels only

Rule Stops when Source
Kneedle At BMI checkpoints it locates the knee of the gain curve, then compares the pre-knee to post-knee slope ratio against 156 - min(relevant at knee, 150). knee_distance='absolute' (default) matches the released implementations; 'signed' is Satopää et al.'s Kneedle. Cormack and Grossman (2016a)
Budget Three quarters of the collection is reviewed, or n >= 10N/r with a knee slope ratio of at least 6. Cormack and Grossman (2016a)
Rule2399 Reviewed documents reach 2399 + 1.2 x relevant found. Cormack and Grossman (2016b)
ReviewHalf Half the known collection is reviewed. Yang and Lewis (2022), TARexp
FixedRound A set number of review batches is done. Yang and Lewis (2022), TARexp
BatchPrecision patience consecutive batches have precision at or below the cutoff. Yang, Lewis, and Frieder (2021)
ConsecutiveIrrelevant A run of count non-relevant documents completes. common heuristic
SAFE Every supplied key paper is found, at least twice the seed positives and min_fraction of the collection are reviewed, and the last consecutive documents are non-relevant. Boetje and van de Schoot (2024)
Oracle The prefix first reaches the target recall, using complete labels. An offline lower bound, never deployable. —

Checkpoint rules — a statistical test at fixed intervals

Rule Stops when Source
CMHHeuristic The biased-urn (BUSCAR) test rejects "recall is still below target" at alpha. Callaghan and Müller-Hansen (2020)
AnytimeCMH The same test with an alpha / (k(k+1)) spending schedule across checkpoints, so the levels sum to alpha. Conditional on the urn p-values being calibrated for the recorded design. original to this package (A. Fletcher)
PoissonPoint (IP-P) A power-law rate fitted to window relevance yields a Poisson upper bound on the documents still unfound, and found >= target x (found + bound); also when the later windows contain nothing relevant. Stevenson and Bin-Hezam (2023)
IPHyperbolic (IP-H) As IP-P with a hyperbolic rate. tail='released' (default) reproduces the released code's expected-remaining formula, which understates it by (1-b)^2; tail='model' integrates the fitted rate. Stevenson and Bin-Hezam (2023)

Control-set rules — an independently screened sample

Rule Stops when Source
QBCB The Jth positive control appears in the ranking, where J is the smallest rank whose binomial bound certifies the target. Lewis, Yang, and Frieder (2021)
TargetRecapture All k = ceil(-ln(1-confidence)/(1-target)) sampled targets have been reviewed. Cormack and Grossman (2016a)
BaselineInclusionRate Relevant documents found reach target x the prevalence estimated from a random pilot. It never fires when the pilot finds nothing. pilot-prevalence heuristic
sample_control Helper: draws an unlabelled worklist to screen independently. Control screening is charged in stop_report and stopping_metrics. —

Estimation rules — recorded probabilities or sampling metadata

Rule Stops when Source
Quant, QuantCI Estimated recall from calibrated probabilities reaches the target; QuantCI first subtracts nstd standard deviations. Pass score_snapshots to replay per-round model scores instead of one fixed column. Yang, Lewis, and Frieder (2021)
AutoStop A Horvitz-Thompson total over a fixed with-replacement sample certifies the target, under loose, strict_v1, or strict_v2; strict_v2 needs collection_size. Not the adaptive procedure. Li and Kanoulas (2020)
SCAL The prefix's Horvitz-Thompson estimate reaches the target share of the estimate over the whole recorded sample. The sample past the cutoff is charged as review cost; pass control_documents() to stopping_metrics so it counts towards recall too. Cormack and Grossman (2016b)
Chao The Chao1 estimate of unseen relevant documents satisfies found >= target x (found + unseen). Simpler than the paper's multi-model estimators. Bron et al. (2025)
BetaBinomial The beta-binomial posterior that few enough relevant documents remain exceeds confidence. Assumes the unreviewed tail is exchangeable with the reviewed prefix, so it errs late on a priority ranking. original to this package (A. Fletcher)
EVPI, EVPIGreedy, EVPISmooth, EVPIBatch The expected value of perfect information about the next document (or window) falls to the screening cost. The document that triggers the stop is included in the reviewed prefix. original to this package (A. Fletcher)

AnytimeCMH, BetaBinomial, and the EVPI family are original to this package rather than replays of published methods, so they carry no parity claim: they are covered by formula tests and the shared leakage and trace checks only.

Learned rule

GRLStop trains a PPO policy over 100 review windows, where the state is the relevance rate in each reviewed window, a logistic-regression prediction for each unreviewed one, the current window, and the target recall. It needs the grl extra. Source: Bin-Hezam and Stevenson (2025).

Datasets and helpers

CLEFTARDataset backs pt.get_dataset('tar:clef2017') and its 2018/2019 siblings: hash-pinned public archives, one candidate pool per topic, complete_ranking() to append unretrieved documents, and label_ranking() to attach qrels only after a ranking exists. Data is downloaded into PyTerrier's cache, never this package.

import pyterrier as pt

dataset = pt.get_dataset('tar:clef2017')

# Replay the pinned, released AutoTAR ranking of one topic.
ranked = dataset.get_results('autotar').query("qid == 'CD008081'")
trajectory = dataset.label_ranking(dataset.complete_ranking(ranked, 'CD008081'))
reviewed = pt.tar.Kneedle().transform(trajectory)

To rank a topic yourself, build an index for it, because each systematic-review topic has its own candidate pool (this needs the tutorial extra):

from tempfile import TemporaryDirectory

import pandas as pd

topic = dataset.get_topics().query("qid == 'CD008081'")
docs = pd.DataFrame(dataset.get_corpus_iter('CD008081'))
with TemporaryDirectory() as index_dir:
    indexref = pt.index.IterDictIndexer(index_dir, meta={'docno': 64}).index(docs.to_dict('records'))
    bm25 = pt.terrier.Retriever(indexref, wmodel='BM25', num_results=len(docs))
    # The topic queries are Boolean PubMed searches; tokenise() reduces them to plain terms.
    ranked = (pt.rewrite.tokenise() >> bm25).transform(topic)
trajectory = dataset.label_ranking(dataset.complete_ranking(ranked, 'CD008081'))

Labels are joined only after a ranking has been created; they must never be made available to a ranker or a live stopping decision.

Metrics

stopping_metrics reports recall, cost, review fraction, reliability, oracle depth, and the CLEF loss_e/loss_r/loss_er components, charging any separately screened controls once. stopping_frontier summarises several targets, and reliability and review_fraction are ir-measures metrics for pt.Experiment.

reliability_summary gives the share of topics reaching the target with a Wilson interval, which is the check for an estimator. coverage_summary gives the share of repeated control draws that certified it, which is the check for a certificate.

Adding a rule

One file per approach, inside the folder for what the rule reads: rules/trajectory/ for labels only, rules/checkpoint/ for a test at fixed checkpoints, rules/control/ for an independently screened sample, and rules/estimation/ for recorded probabilities or sampling metadata.

# src/pyterrier_tar/rules/trajectory/my_rule.py
from .._base import _TrajectoryStoppingRule, _batch_positions


class MyRule(_TrajectoryStoppingRule):
    """Stops once a batch holds no relevant documents."""

    _equation = 'fire if the batch ending at n has no relevant document'

    def __init__(self, batch_size: int = 200, initial_documents: int = 1):
        if batch_size < 1 or initial_documents < 1:
            raise ValueError('batch sizes must be positive')
        self.batch_size = batch_size
        self.initial_documents = initial_documents

    def _checkpoint_rows(self, ranked, labels):
        for stop in _batch_positions(len(labels), self.batch_size, self.initial_documents):
            batch = labels[max(0, stop - self.batch_size):stop]
            yield {'index': stop, 'relevant_in_batch': int((batch > 0).sum()),
                   'fired': not (batch > 0).any()}

_checkpoint_rows yields one dict per checkpoint, with at least index and fired; the base class turns it into the stop, stop_report, stop_trace, and calculation_trace, so those cannot disagree. A rule that is not checkpoint-shaped implements _stop_position(labels) instead, and one that needs other columns implements _stop_position_for_results(ranked, labels). Add _trace_details for the columns stop_trace should carry.

The row for checkpoint k must use only labels[:k]. Export the class from the group's __init__.py, from rules/__init__.py, and from pyterrier_tar/__init__.py; the shared tests then pick it up, including the leakage check and the report/trace consistency check. New rules are expected to come with a test that re-derives the stop from the rule's own definition, as tests/test_rule_validity.py does for the others.

Scope and safety boundary

pyterrier-tar replays recorded, labelled trajectories. It is not a live active-learning or screening controller, and it does not certify target recall in deployment. Labels beyond the reviewed prefix must not be exposed to a stopping rule. Separately sampled controls must be independently screened and accounted for in review cost.

Install

python -m pip install pyterrier-tar
python -m pip install 'pyterrier-tar[grl]'

The grl extra installs gymnasium, scikit-learn, and stable-baselines3 for GRLStop; plot adds matplotlib for the figures; tutorial adds PyTerrier's Java support and matplotlib for the notebooks that build Terrier indexes.

The documents this README names, the example scripts, the notebooks, and the tests are in the source distribution, not the wheel:

python -m pip download pyterrier-tar --no-deps --no-binary :all:
tar -xzf pyterrier_tar-*.tar.gz

The unpacked directory holds BENCHMARKING.md, EVALUATION.md, PUBLISHED_EVIDENCE.md, GAPS.md, RESULTS.md, RESULTS_UNCERTAINTY.md, CHANGELOG.md, examples/, and tests/. The same archive is on the PyPI files page.

Quick start

import numpy as np
import pandas as pd
from pyterrier import tar

# A recorded review: ranked documents, with the label revealed as each is read.
# Relevance thins out down the ranking, as it does in a real screening run.
rng = np.random.default_rng(0)
size = 6000
prevalence = .7 * np.exp(-np.arange(size) / 300) + .002
trajectory = pd.DataFrame({
    'qid': 'CD008081',
    'docno': [f'd{position}' for position in range(size)],
    'rank': range(size),
    'label': (rng.random(size) < prevalence).astype(int),   # 1 for relevant
})

rule = tar.Kneedle()
reviewed = rule.transform(trajectory)          # the prefix the rule would have read
report = rule.stop_report(trajectory)          # stop, fired, review cost
print(rule.stop_trace(trajectory).iloc[0].reason)
print(rule.calculation_trace(trajectory, index=int(report.stop.iloc[0])).tail())

qrels = trajectory.loc[trajectory.label > 0, ['qid', 'docno', 'label']]
print(tar.stopping_metrics(trajectory, reviewed, qrels, target_recall=.95, report=report))

On this trajectory Kneedle stops after 2,001 of 6,000 documents, having found 94.7% of the relevant ones; stop_trace says the knee slope ratio reached its dynamic threshold, and calculation_trace shows the ratio at every checkpoint it tested.

Benchmarking

The experiment API runs recorded histories, saves per-topic results, and reports uncertainty and failures. The input ranking is fixed before labels are attached; this workflow does not train a ranker or select a stopping rule.

from pathlib import Path
import pyterrier_tar as tar

# trajectory has qid, docno, rank, label and covers the full collection.
data = tar.ReplayDataset("my-collection", "recorded-ranker", trajectory)
rules = [
    tar.RuleSpec("IP-P", lambda c: tar.PoissonPoint(c.target, c.collection_size)),
    tar.RuleSpec("50 irrelevant", lambda c: tar.ConsecutiveIrrelevant(50),
                 config={"count": 50}),
]
output = Path("artifacts/my-experiment")
results = tar.run_experiment([data], rules, targets=[.8, .9, .95],
                             seeds=[0], output=output)
summary = tar.experiment_summary(results)
summary.to_csv(output / "summary.csv", index=False)
(output / "REPORT.md").write_text(tar.experiment_report(results))
  • Rules. A factory receives RuleContext(trajectory, target, seed, collection_size) and returns a fresh rule. Built-in rules are classified from the catalogue; a custom class needs promise= (heuristic, estimator, certificate, or offline bound), which is metadata, not a verified guarantee. A factory that samples must draw from context.rng(), which is keyed by topic, target and seed; default_rng(context.seed) alone repeats one stream for every topic and target.
  • Repetition. repetitions="once" uses the first seed and "seeds" uses every seed. Certificates default to "seeds" and other rules to "once".
  • Resume and provenance. Completed tasks are written atomically under tasks/, and repeating the call resumes them. manifest.json binds the parameters, input-frame hashes, package source hash and dependency versions; changing the manifest requires a new directory. Only one writer may use a directory.
  • Partial histories. Omitting qrels declares the trajectory to be the entire collection, and the runner cannot infer unseen relevant documents. For a partial history, pass full-pool qrels including negatives.
  • Uncertainty and failures. experiment_summary gives target attainment with Wilson 95% intervals, mean review fraction with bootstrap intervals, minimum recall, recall shortfalls, and counts of firings, full-review fallbacks, exhausted partial histories and zero-relevant topics. worst_topics keeps the topic, target and seed of individual failures. For a certificate, each row is one topic across independent control draws, and certification frequency is separate from reaching the target: reviewing everything can reach the target without producing a certificate. The intervals are descriptive and do not adjust for selecting among methods or targets.
  • Baseline rankings. bm25_ranking and random_ranking are generated before qrels are joined, and complete_rankings(pool, ranking) appends unranked documents, so rankers are compared on the same pool. bm25_ranking is a specified in-memory baseline and claims no Terrier scoring parity.
  • Review histories. read_review_history("events.csv", scores="scores.csv") reads portable CSV events (qid,docno,rank,label,round); import_tarexp and read_tarexp read TARexp ledgers. Native TARexp checkpoints are gzipped Python pickles and may execute code, so read_tarexp requires trusted=True: only load checkpoints you trust. TARexp records batch membership, not within-batch review order, so a document-level stop inside an imported batch is not a reconstructed historical decision.
  • Figures. With the plot extra, figure_recall_effort, figure_target_misses, figure_control_cost, figure_stopping_outcomes and figure_stopping_diagnostic return figures 180 mm wide, and save_figure(fig, "figures/name") writes PDF, SVG and 600-dpi PNG.

BENCHMARKING.md has the rest: what the manifest records, ranking comparisons on CLEF with examples/benchmark_rankings.py, the history formats, and what each figure shows.

Development

From the unpacked source distribution:

python -m pip install -e '.[test,grl,tutorial]'
pytest

test runs the suite, grl adds the GRLStop tests, and tutorial the notebooks. On a CUDA-free machine, UV_TORCH_BACKEND=cpu keeps grl from pulling the GPU wheels.

Results on CLEF

Every rule replays the released AutoTAR ranking of each topic in CLEF 2017, 2018, and 2019 (91 topics with a relevant document: 30 from 2017, 30 from 2018, 31 from 2019), completed with the pool's unranked documents. Review is the share of each topic's pool screened, averaged over topics; for certificates it includes the control sample.

These numbers describe one ranker. A rule that reads the ranking, which is every rule here except the certificates, can do much worse when the ranking is weaker. Rules are grouped by what they promise, so compare within a group, not across groups.

An estimator promises the target if its model holds. Reliability is the share of topics that reached the target; the rule was asked for 95% confidence, so 95% reliability is what it promises. The Oracle reads every label, so no deployable rule can match it; it marks the least review each target could have needed.

Rule Reliability at 0.8 / 0.9 / 0.95 Review at 0.8 / 0.9 / 0.95
CMHHeuristic 100.0% / 100.0% / 100.0% 47.6% / 60.8% / 72.1%
AnytimeCMH 100.0% / 100.0% / 100.0% 63.8% / 78.1% / 90.1%
PoissonPoint (IP-P) 100.0% / 100.0% / 100.0% 27.7% / 28.6% / 28.9%
IPHyperbolic (IP-H) 93.4% / 90.1% / 89.0% 16.1% / 16.8% / 18.1%
IPHyperbolic (IP-H), tail='model' 100.0% / 97.8% / 97.8% 17.6% / 18.5% / 20.0%
BetaBinomial 100.0% / 100.0% / 100.0% 90.3% / 96.5% / 99.0%
Oracle (offline bound) 100.0% / 100.0% / 100.0% 5.1% / 6.5% / 7.8%

A heuristic takes no target and promises nothing about recall. It stops in the same place whatever the target; the columns show how often that happened to reach each one.

Rule Review Reached 0.8 / 0.9 / 0.95
Kneedle 54.7% 100.0% / 100.0% / 100.0%
Budget 43.8% 100.0% / 100.0% / 100.0%
Rule2399 72.3% 100.0% / 100.0% / 100.0%
ReviewHalf 58.8% 100.0% / 100.0% / 100.0%
BatchPrecision 31.5% 100.0% / 98.9% / 96.7%
ConsecutiveIrrelevant 20.8% 100.0% / 97.8% / 95.6%
SAFE (no seed positives) 18.0% 100.0% / 100.0% / 94.5%

A certificate promises the target with 95% probability over the draw of its control sample, so it is measured by coverage: the share of 20 simulated samples per topic that reached the target. A topic with too few relevant documents cannot be certified, and the rule then reviews everything.

Rule Target Coverage Topics it could certify Review (of which control)
QBCB 0.8 98.0% 71 of 91 48.5% (42.3%)
QBCB 0.9 98.2% 60 of 91 65.6% (60.0%)
QBCB 0.95 99.1% 39 of 91 83.5% (79.0%)
TargetRecapture 0.8 98.5% 71 of 91 50.0% (43.8%)
TargetRecapture 0.9 98.6% 59 of 91 66.5% (60.9%)
TargetRecapture 0.95 99.2% 38 of 91 83.8% (79.4%)

Quant, QuantCI, AutoStop, SCAL, Chao, the EVPI family, GRLStop, FixedRound, and BaselineInclusionRate are not in these tables: each needs an input the CLEF release does not contain, a trained policy, or a parameter whose value would be an arbitrary choice.

These tables are copied from RESULTS.md, which adds the breakdown by CLEF year. RESULTS_UNCERTAINTY.md gives the intervals behind each row and the largest individual shortfalls. examples/results_table.py regenerates both.

Published Evidence

A passing check does not turn a heuristic rule into a deployment recall certificate. Each row says what was compared and how far the claim goes.

Rule or method Checked against What the check establishes
stopping_metrics the official CLEF tar_eval.py on its two sample runs, all 20 topics recall, total cost, and the three loss fields match at 3 decimals; no claim for the CLEF metrics it does not expose, such as NCG, AP, or WSS
IPHyperbolic (IP-H) the authors' released rankings and results for CLEF 2017--19 strict parity: recall, cost, reliability, loss_er, and relative error in all nine settings at 3 decimals
PoissonPoint (IP-P) the same release, CLEF 2017 strict parity: every published IP-P and IP-H aggregate at 3 decimals
Oracle the released point-process baseline table all nine CLEF 2017--19 oracle rows match
Kneedle Sneyd and Stevenson's source on the CLEF 2017 Waterloo run the source's total effort (86,243) and 0.998333 mean recall, with min_documents=1, batch_size=1; the default fixed-200 replay is a compatibility mapping
CMHHeuristic buscarpy, pinned 216 p-values and five retrospective stop positions match; source-definition parity, not a paper reproduction
QBCB Lewis, Yang, and Frieder's Table 1 every published r -> J control rank; no numerical paper parity, because the paper's inputs were not released
AutoStop the released AutoStop source, CLEF topic CD008081 the loose, strict_v1, and strict_v2 estimators on a fixed sampling distribution; not the adaptive procedure
Quant, QuantCI, BatchPrecision, Rule2399, ReviewHalf, FixedRound TARexp's definitions definitions follow TARexp and Quant's variance mirrors it; no paper result is asserted
Budget, SCAL, Chao, TargetRecapture, BaselineInclusionRate their published definitions definition checks only; Chao is the classic Chao1, simpler than the paper's estimators
GRLStop the released GRLStop source compared line by line; no trained policies or result tables were released, so there is nothing numerical to reproduce
AnytimeCMH, BetaBinomial, the EVPI family nothing external original to this package: formula tests and the shared leakage and trace checks only

Every rule that decides from a reviewed prefix is also checked for leakage: changing the labels after its stop, or after a trace index, never changes it.

PUBLISHED_EVIDENCE.md is the full ledger, with the pinned revisions and hashes behind each row. The examples under examples/ are intentionally separate from CI when they require network access, large public archives, or locally held trajectories.

For real-data or external-source smoke checks:

python examples/published_cmh_buscar.py
python examples/published_clef_tar_eval_metrics.py
python examples/published_ip_h_clef.py
python examples/published_kneedle_clef2017.py
python examples/published_point_process_clef2017.py
python examples/published_autostop_clef2017.py
PYTHONPATH=examples python examples/published_baselines_clef.py

Run them from a scratch directory: each caches its inputs in data/ relative to the working directory, about 1 GB in total.

examples/notebooks/parity_published.ipynb runs all of them and shows one parity table. Every notebook is either a demo_ (how the package behaves) or a parity_ (a published result reproduced):

Notebook Purpose
demo_stopping_rules Every public stopping rule on a small synthetic trajectory: its decision, trace, and replay-only boundary.
demo_trajectory_rules Why each trajectory-only rule fires, drawn: revealed labels, checkpoints, the trigger quantity, and the stop.
demo_workflows Two worked workflows on real collections: a CLEF topic from index to stop, and control-set stopping on Vaswani with its screening cost.
demo_fresh_rankers_stopping Build fresh BM25/TF-IDF rankings, run editable stopping methods in memory, and display publication figures.
demo_multitopic_rankers_stopping The same across ten topics by default or a whole CLEF year, with all-topic outcomes and a paired diagnostic.
demo_publication_figures Load and verify saved benchmark results, then display and export the publication figures.
parity_published One table for every published claim: runs each networked parity script and asserts its result.
parity_archives Four audits of other groups' released results: Repke et al. (2026), König et al. (2024), Chao et al. (2024), and the SYNERGY v2 Kneedle screen. Each needs archives you download, and skips when they are absent.

published_kneedle_clef2017.py checks Kneedle's released 86,243 effort. The same seven scripts run weekly in CI's published-parity job. The notebooks that build Terrier indexes need the tutorial extra, which installs PyTerrier's Java support.

Only the metrics asserted by each evidence check should be treated as validated. For example, the CLEF metric checker asserts only the official recall, total cost, and loss fields exposed by stopping_metrics(), while the IP-H CLEF checker asserts recall, cost, reliability, loss, and relative-error parity across all nine CLEF 2017--19 settings.

Known gaps

What this package does not implement or cannot reproduce, and why. It replays recorded reviews, so nothing here establishes how a rule behaves when it steers a live one.

These published results need data the authors did not release. Nothing in the package should claim them.

Result Missing input
QBCB's RCV1 figures the 20% subset, seeds, active-learning trajectories, and control-sample identities
Yang et al. (2021) heuristics on RCV1 the saved review histories
Adaptive AutoStop the per-draw sampling distributions, which a trajectory cannot reconstruct
Chao et al. Table 7 integer displays the analysis and formatting code; raw aggregates do not round to the printed integers (57.566 prints as 57)
The Knee rows of the point-process baseline table the 2018 qrels are rebuilt in that repository from a PID list it does not ship; published_baselines_clef.py reports the difference instead of asserting it
TM and TM-adapted rows of the same table the released target method draws its target set with an unseeded random.choice
RLStop and GRLStop paper results no trained policies or result tables were released

These methods are named but not implemented.

Name State
Score-distribution stopping (Hollmann and Eickhoff) not implemented; QuantCI over calibrated probabilities is the nearest rule
The adapted target method of the point-process papers not implemented; TargetRecapture is Cormack and Grossman's original target method
TARexp-style Knee and Budget TARexp takes the maximum slope ratio over every round split, measured in rounds. This package follows Cormack and Grossman's knee. An opt-in mode would be needed for TARexp-equivalent numbers.
Adaptive AutoStop, a live S-CAL controller out of scope: this package replays recorded trajectories and does not select documents

And these are limits of the engineering.

Gap Notes
CLEF archives come from one mirror npai.science.uu.nl, hash-pinned. The hashes protect the content, not its availability
9 of 119 SYNERGY datasets will not compose missing source files or server errors at the pinned commit
Parity runs weekly, archive audits by hand the published-parity job downloads about 1 GB; the archive notebooks need artifacts you supply

These tables are copied from GAPS.md. New findings about limits belong there rather than in a commit message.

Compatibility policy

The supported PyTerrier range is declared in pyproject.toml. New code must import only public pyterrier names; imports from private modules such as pyterrier._ops are prohibited. Each supported PyTerrier release will receive a focused compatibility test before its range is widened.

License and notices

This source is licensed under the Mozilla Public License 2.0. LICENSE.txt and THIRD_PARTY_NOTICES.md are installed with the package, in the licenses folder of its dist-info directory.

The package contains no third-party source code. It was first written in a fork of PyTerrier and extracted into this package; it uses PyTerrier only through PyTerrier's public API. The stopping rules implement published methods from their papers. Where a rule reproduces a released configuration, such as IPHyperbolic's released tail formula, it is an implementation of the published method, not a copy of the authors' source code.

Metadata

Release files for pyterrier-tar 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyterrier-tar 0.1.1
File Size Uploaded
pyterrier_tar-0.1.1.tar.gz 220.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyterrier-tar 0.1.1
File Interpreter ABI Platform
pyterrier_tar-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 327.0 kB

Release files / pyterrier_tar-0.1.1.tar.gz

Download URL pyterrier_tar-0.1.1.tar.gz
Size 220.0 kB
Tags Source
SHA-256 checksum
How to use checksums
e7242a730f5894dcc8bae5bb80d6a9488f5c1031b0bb359d62d72b2c65bdb172
BLAKE2b-256 checksum
How to use checksums
388e49532406043b27280445ec187eba28b2f65ecb6a8b5285d9a6d7966ee55f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release files / pyterrier_tar-0.1.1-py3-none-any.whl

Download URL pyterrier_tar-0.1.1-py3-none-any.whl
Size 107.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6eca535a59a2cc19b73ed4ad0465a4df381560e5c010542759db68c6cfdce4df
BLAKE2b-256 checksum
How to use checksums
fbd4d47567abb2b9925f18844b6177285ec6ffb2cf05692869b48bbac8ef2833
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page