szl-crosscheck
Two independent implementations of the same measurement disagree quietly more
often than teams admit. This package ends that: feed it the receipt chains from
both harnesses and get one verdict — CONSISTENT, DIVERGENT (with the
metric and max relative delta named), INCOMPARABLE (no shared MEASURED lanes),
or INVALID (a chain failed verification — fail closed).
Built to settle the duplicated bench layer in this org: the stdlib repos
(szl-retrieval-bench, szl-engine-bench, szl-quant-bench) and the FastAPI
planes (retrieval-bench, frontier-bench, quant-curve). Their receipt
formats differ: the retrieval HTTP runner returns an unchained RunReceipt,
while the stdlib runner uses a self_hash chain. Explicit adapters preserve
these native records inside an adapter-owned integrity envelope. This does
not claim that the HTTP service emitted a chain or attested the caller's context.
Doctrine
- Both chains verify before anything is compared. Tampering voids the run.
- Measured results require matching machine identity, CPU, OS, Python version, dataset/model revision, parameters, and corpus/query/qrels hashes. Source repository commits are recorded separately so independent implementations can differ. These fields are declared evidence, not hardware attestation.
- Only shared MEASURED lanes with shared metric keys are compared.
- Divergence is always named: lane, metric, max relative delta.
- The dual receipt commits to both terminal hashes and every lane verdict — same chains, same tolerance, same receipt, any machine.
Usage
pip install -e . pytest && python -m pytest tests/ -q
import json
from szl_crosscheck import crosscheck
a = json.loads(open("stdlib-chain.json").read())
b = json.loads(open("plane-chain.json").read())
print(crosscheck(a, b, rel_tol=0.01)["verdict"])
Scope
Python 3.11+, standard library only. Never re-measures; verifies and compares what the harnesses already receipted. Numbers from different hardware remain incomparable by design.
Native BM25 adapters
szl_crosscheck.adapters.stdlib_retrieval(chain, context) validates the native
stdlib chain and requires its final run.context to equal the supplied context.
Its final run.result must be a measured BM25 result at the declared cutoff.
fastapi_retrieval(response, context, captured_inputs) checks the HTTP response's
dataset, qrels, model/config and result hashes, then binds all three captured
input hashes. When query_hash is supplied, it must also match the captured
queries using the native producer's SHA-256 over UTF-8 JSON with sorted keys,
default ASCII escaping and default (noncompact) separators. Null, malformed or
mismatching supplied hashes are rejected. This matches the
pinned retrieval-bench producer.
The normalized envelope records native_query_hash_state = MATCHED when that
hash matches. Historical responses that omit the field remain supported with
native_query_hash_state = UNVERIFIED_LEGACY and an explicit capture-only limitation;
an absent field is different from a supplied null or invalid hash. In both cases
request capture and context remain caller-declared. Matching hashes establish
input integrity, not proof of server execution, semantic accuracy or signer identity.
Both adapters compare only the shared nDCG@10 and Recall@10 by default; they
retain the complete native records, including non-compared metrics.
Context is an object containing machine (identity, cpu, os, python),
dataset_revision, model_revision, parameters (top_k, k1, b),
input_hashes (full SHA-256 values for corpus, queries, qrels) and
source (repository and full 40-character commit). Input hashes use UTF-8
JSON with sorted keys, compact separators, default ASCII escaping and no NaN.
Do not relabel fixture inputs as a real benchmark. Source commits and input
hashes support reproducibility; unsigned receipts do not prove signer identity.
Recorded native run
evidence/2026-09-05/crosscheck-native-scifact-20260905.json records an actual
same-machine run over all 5,183 cached public SciFact documents and 300 test
queries. The FastAPI side was called using an actual loopback HTTP POST and
the ephemeral server was stopped afterward. With identical inputs and BM25
parameters, nDCG@10 was 0.6380 (stdlib) and 0.6643898053153046 (FastAPI).
The 3.972% relative difference is DIVERGENT at the declared 1% tolerance.
The native tokenizers differ; this receipt demonstrates the disagreement,
not interchangeability or a performance ranking. Full normalized input chains,
native records, source pins and input hashes are retained. CI recomputes both
the complete bundle hash and dual report; it does not rerun the corpus benchmark.
Independent Lean-build comparison
schemas/lean-build-comparison-v1.json and
szl_crosscheck.validate_lean_build_receipt define a second, deliberately
narrow comparison plane for large Lean builds. The target is bound to an exact
artifact digest, Lean version, entrypoint, theorem name, reference-statement
digest and declared axiom surface. Each harness must record its own implementation
digest, runner identity, exact source revision, rebuild and kernel outcomes,
statement observation, axiom observation, final-chain sorry count, dependency
cone and log digest.
The semantic validator recomputes every comparison field and the final verdict:
CONSISTENTmeans all required bounded observations agree across independently identified harness implementations and runners;DIVERGENTnames a complete but disagreeing statement, axiom, dependency or kernel observation;INCOMPARABLEis mandatory when independence, target identity, execution, statement fidelity, dependency evidence or final-chain evidence is incomplete.
A caller cannot write CONSISTENT over disagreeing or incomplete evidence: the
receipt fails validation unless its comparison, verdict and ordered reasons are
exactly derivable from the observations. This contract never emits BREAK or
PROOF_VALID. Comparison consistency is not proof validity, theorem correctness,
security, or production readiness; the Lean kernel's own check remains the
relevant formal check.
from szl_crosscheck import validate_lean_build_receipt
validated = validate_lean_build_receipt(receipt)
print(validated["verdict"], validated["verdict_reasons"])
License
Apache-2.0 — see LICENSE (same text as the org's other bench repos).
Metadata
Release files for szl-crosscheck 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| szl_crosscheck-0.3.0.tar.gz | 23.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| szl_crosscheck-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 41.2 kB
Release files / szl_crosscheck-0.3.0.tar.gz
| Download URL | szl_crosscheck-0.3.0.tar.gz |
|---|---|
| Size | 23.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
db628d51afd56b30fea21bfbbb4d71cc6211c9526c4b4682014c1d84c44f1aa8
|
|
BLAKE2b-256 checksum How to use checksums |
67b14fc463580d958907045c3b2635a3792109cc59b77921bf59ca5ab137a398
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / szl_crosscheck-0.3.0-py3-none-any.whl
| Download URL | szl_crosscheck-0.3.0-py3-none-any.whl |
|---|---|
| Size | 17.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e691de96cf186f80bd2969b1b8e6dcc8dce34b3db75c264b6ab0dbbda1ea0086
|
|
BLAKE2b-256 checksum How to use checksums |
136d0782e68cb861b03261900bfed7ff37c21bfa5fcc4a49fc9a26fd17cedf85
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log