Skip to main content

szl-retrieval-bench

Honest retrieval benchmark lane for the SZL estate. Sparse BM25 baseline, classical TF-IDF dense lane, RRF hybrid fusion, and ranking metrics (nDCG@k, Recall@k, P@k, R-precision, MRR, MAP) with fairness gates and hash-chained receipts.

Doctrine: a benchmark that cannot run returns BLOCKED with a reason. It never fabricates a metric. Run states: MEASURED | BLOCKED | INVALID | FAILED.

Install and verify

pip install -e . pytest
python -m pytest tests/ -q                  # full offline suite
python -m szl_retrieval_bench.harness       # JSON demo; P@10 + R-precision in every lane/receipt

Metric API

from szl_retrieval_bench.metrics import precision_at_k, r_precision

ranked = ["d1", "d2", "d3", "d4", "d5"]
relevant = {"d1", "d3", "d9"}

assert precision_at_k(ranked, relevant, 2) == {
    "state": "MEASURED", "P@2": 0.5,
}
assert r_precision(ranked, relevant) == {
    "state": "MEASURED", "R_precision": 0.6667, "R": 3,
}

P@k always divides by k, including when fewer than k results were returned. k < 1 or a non-integer cutoff is INVALID; R-precision is INVALID when there are no relevant documents. A valid ranking with no overlap is MEASURED at 0.0 and remains present in per-query output, aggregates, and receipts.

Lanes

Synthetic answer fixtures

python -m szl_retrieval_bench.fixture scores only caller-supplied synthetic question/answer fixtures with fixture_exact_match_v1: NFC normalization, casefolding, collapsed Unicode whitespace, and strict equality retaining punctuation and signs. It does not run retrieval, generation, or a provider. External performance remains UNMEASURED; this is not an official LongMemEval, LoCoMo, or STATE-Bench evaluation. The existing ranking APIs are unchanged.

Use a JSON array of {question_id, question, answer} objects and JSONL {question_id, hypothesis} predictions, with unique nonempty string IDs:

python -m szl_retrieval_bench.fixture --dataset fixture.questions.json \
  --dataset-variant fixture --predictions fixture.predictions.jsonl \
  --system synthetic --judge fixture_exact_match_v1 --out fixture-receipts

Missing predictions count as incorrect; empty datasets, empty prediction sets, all-empty hypotheses, duplicate/unknown IDs, malformed JSON, and unsupported judge/dataset modes raise errors. Inputs are capped at 1 MiB and 1,000 rows. The legacy exact option aliases strict fixture equality, never substring matching.

Each unique output directory retains input/source snapshots, verdicts, and an unsigned hash-bound receipt. fixture.verify_receipt(path, expected_evaluator_sha256=trusted_digest) recomputes consistency and optionally checks an independently obtained source digest. It never executes the snapshot. Unsigned coordinated rewrites are not authenticated evidence. Normal imports honestly label source identity as an unverified file snapshot. Output and input parents must be private, caller-owned local directories: portable link checks do not defeat concurrent hostile directory replacement. See integration provenance for coverage and limits.

Retrieval ranking

  • bm25 - stdlib BM25 (k1=1.5, b=0.75). No downloads, no network.
  • tfidf-dense — classical dense retrieval: L2-normalized TF-IDF vectors, cosine similarity. Real dense-vector math, deterministic, honestly labeled classical (not a neural embedding). A neural adapter plugs into the same rank(query, doc_ids) interface without touching the harness.
  • hybrid_rrf — RRF(k=60) fusion of BM25 and a pluggable dense ranker. Calling it without a dense ranker returns BLOCKED, by design.
  • compare — fairness gate: runs covering different query sets are INVALID.
  • Receipts — every comparison can emit a SHA-256 hash-chained UNSIGNED_HONEST receipt. The receipt hashes the source harness declaration szl-retrieval-bench so downstream aggregators can reject relabelling. This is a hashed source declaration, not issuer authentication: the chain proves integrity + order of its contents, not who produced them.

Multi-vector / late-interaction (ColBERT-style) lives in a separate lane: different memory profile, different fairness constraints. Not mixed here.

Changelog highlights

  • current: comparison receipts bind the szl-retrieval-bench harness name in the hashed payload so Wave 1 consolidation can verify source declaration without inferring identity from a filename or caller-supplied label.
  • v0.3.0: P@k and R-precision added to the metric API, all measured lanes, demo output, aggregate leaderboards, and hash-chained comparison receipts; invalid cutoffs and undefined R-precision fail closed.
  • v0.2.0: dense lane added; fixed a latent circular-reference bug in receipted comparisons (found by the new three-lane test before push — receipts now embed a snapshot, never a self-reference).
  • v0.1.0: BM25, RRF, metrics, fairness gates, receipts, CI.

Apache-2.0 · Doctrine v11 · SZL Holdings

Metadata

Release files for szl-retrieval-bench 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for szl-retrieval-bench 0.3.0
File Size Uploaded
szl_retrieval_bench-0.3.0.tar.gz 26.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for szl-retrieval-bench 0.3.0
File Interpreter ABI Platform
szl_retrieval_bench-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 46.9 kB

Release files / szl_retrieval_bench-0.3.0.tar.gz

Download URL szl_retrieval_bench-0.3.0.tar.gz
Size 26.6 kB
Tags Source
SHA-256 checksum
How to use checksums
26734c4d4d48d89e040eff8bccf8e65a2035578a1755d70b8bfa3808f0a82ff9
BLAKE2b-256 checksum
How to use checksums
ba8bf7d07f9879a0e77d4f3463f055c16c36c12ed69effe8500996f4f9a61092
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / szl_retrieval_bench-0.3.0-py3-none-any.whl

Download URL szl_retrieval_bench-0.3.0-py3-none-any.whl
Size 20.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
823395093f7d58d9b51d9658476d71d0406b87b4a66f59e0ebab03e393453660
BLAKE2b-256 checksum
How to use checksums
cd88503c36865464f1ce530afe109fdd0184bbf1c4e659b5ae9f2de744bbafbf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page