szl-retrieval-bench
Honest retrieval benchmark lane for the SZL estate. Sparse BM25 baseline, classical TF-IDF dense lane, RRF hybrid fusion, and ranking metrics (nDCG@k, Recall@k, P@k, R-precision, MRR, MAP) with fairness gates and hash-chained receipts.
Doctrine: a benchmark that cannot run returns BLOCKED with a reason.
It never fabricates a metric. Run states: MEASURED | BLOCKED | INVALID | FAILED.
Install and verify
pip install -e . pytest
python -m pytest tests/ -q # full offline suite
python -m szl_retrieval_bench.harness # JSON demo; P@10 + R-precision in every lane/receipt
Metric API
from szl_retrieval_bench.metrics import precision_at_k, r_precision
ranked = ["d1", "d2", "d3", "d4", "d5"]
relevant = {"d1", "d3", "d9"}
assert precision_at_k(ranked, relevant, 2) == {
"state": "MEASURED", "P@2": 0.5,
}
assert r_precision(ranked, relevant) == {
"state": "MEASURED", "R_precision": 0.6667, "R": 3,
}
P@k always divides by k, including when fewer than k results were
returned. k < 1 or a non-integer cutoff is INVALID; R-precision is
INVALID when there are no relevant documents. A valid ranking with no
overlap is MEASURED at 0.0 and remains present in per-query output,
aggregates, and receipts.
Lanes
Synthetic answer fixtures
python -m szl_retrieval_bench.fixture scores only caller-supplied synthetic
question/answer fixtures with fixture_exact_match_v1: NFC normalization,
casefolding, collapsed Unicode whitespace, and strict equality retaining
punctuation and signs. It does not run retrieval, generation, or a provider.
External performance remains UNMEASURED; this is not an official LongMemEval,
LoCoMo, or STATE-Bench evaluation. The existing ranking APIs are unchanged.
Use a JSON array of {question_id, question, answer} objects and JSONL
{question_id, hypothesis} predictions, with unique nonempty string IDs:
python -m szl_retrieval_bench.fixture --dataset fixture.questions.json \
--dataset-variant fixture --predictions fixture.predictions.jsonl \
--system synthetic --judge fixture_exact_match_v1 --out fixture-receipts
Missing predictions count as incorrect; empty datasets, empty prediction sets,
all-empty hypotheses, duplicate/unknown IDs, malformed JSON, and unsupported
judge/dataset modes raise errors. Inputs are capped at 1 MiB and 1,000 rows.
The legacy exact option aliases strict fixture equality, never substring matching.
Each unique output directory retains input/source snapshots, verdicts, and an
unsigned hash-bound receipt. fixture.verify_receipt(path, expected_evaluator_sha256=trusted_digest) recomputes consistency and optionally
checks an independently obtained source digest. It never executes the snapshot.
Unsigned coordinated rewrites are not authenticated evidence. Normal imports
honestly label source identity as an unverified file snapshot. Output and input
parents must be private, caller-owned local directories: portable link checks
do not defeat concurrent hostile directory replacement. See
integration provenance for coverage and limits.
Retrieval ranking
bm25- stdlib BM25 (k1=1.5, b=0.75). No downloads, no network.tfidf-dense— classical dense retrieval: L2-normalized TF-IDF vectors, cosine similarity. Real dense-vector math, deterministic, honestly labeled classical (not a neural embedding). A neural adapter plugs into the samerank(query, doc_ids)interface without touching the harness.hybrid_rrf— RRF(k=60) fusion of BM25 and a pluggable dense ranker. Calling it without a dense ranker returnsBLOCKED, by design.compare— fairness gate: runs covering different query sets areINVALID.- Receipts — every comparison can emit a SHA-256 hash-chained
UNSIGNED_HONESTreceipt. The receipt hashes the source harness declarationszl-retrieval-benchso downstream aggregators can reject relabelling. This is a hashed source declaration, not issuer authentication: the chain proves integrity + order of its contents, not who produced them.
Multi-vector / late-interaction (ColBERT-style) lives in a separate lane: different memory profile, different fairness constraints. Not mixed here.
Changelog highlights
- current: comparison receipts bind the
szl-retrieval-benchharness name in the hashed payload so Wave 1 consolidation can verify source declaration without inferring identity from a filename or caller-supplied label. - v0.3.0: P@k and R-precision added to the metric API, all measured lanes, demo output, aggregate leaderboards, and hash-chained comparison receipts; invalid cutoffs and undefined R-precision fail closed.
- v0.2.0: dense lane added; fixed a latent circular-reference bug in receipted comparisons (found by the new three-lane test before push — receipts now embed a snapshot, never a self-reference).
- v0.1.0: BM25, RRF, metrics, fairness gates, receipts, CI.
Apache-2.0 · Doctrine v11 · SZL Holdings
Metadata
Release files for szl-retrieval-bench 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| szl_retrieval_bench-0.3.0.tar.gz | 26.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| szl_retrieval_bench-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 46.9 kB
Release files / szl_retrieval_bench-0.3.0.tar.gz
| Download URL | szl_retrieval_bench-0.3.0.tar.gz |
|---|---|
| Size | 26.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
26734c4d4d48d89e040eff8bccf8e65a2035578a1755d70b8bfa3808f0a82ff9
|
|
BLAKE2b-256 checksum How to use checksums |
ba8bf7d07f9879a0e77d4f3463f055c16c36c12ed69effe8500996f4f9a61092
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / szl_retrieval_bench-0.3.0-py3-none-any.whl
| Download URL | szl_retrieval_bench-0.3.0-py3-none-any.whl |
|---|---|
| Size | 20.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
823395093f7d58d9b51d9658476d71d0406b87b4a66f59e0ebab03e393453660
|
|
BLAKE2b-256 checksum How to use checksums |
cd88503c36865464f1ce530afe109fdd0184bbf1c4e659b5ae9f2de744bbafbf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log