Skip to main content

crucible: A judgment engine: register a thesis, steelman each claim, measure against a substrate, refine the weakest axis.

A judgment engine: register a thesis, steelman each claim, measure against a substrate, refine the weakest axis.

PyPI license: crucible Fair-Source CI downloads python: 3.11+ deps: none (core)

crucible turns a thesis into a set of claims, each paired with the observation that would refute it. Independent adversaries steelman every claim by proposing the strongest test, the engine measures each one against a substrate oracle, and the weakest axis gets refined across rounds: strengthen the substrate, sharpen the measurement, or amend the thesis. The result is a verdict per claim, MATCH, DRIFT, or UNVERIFIABLE, grounded in the measurement rather than a judge's opinion. Every run writes a record you can re-check.

Project Telos | gather | crucible | index | forum | telos | learn | emet | buildlang

Highlights

  • Verdicts that recompute from the record. verdict_for(claim, measurement) is a pure function: within tolerance is MATCH, outside is DRIFT, absent or unmeasurable is UNVERIFIABLE, fail-closed. No model sits in the verdict step, so a fluent assertion cannot move a rechecked result.
  • One-command runs with cleanroom review packets. crucible run --bundle DIR executes steelman, measurement, witnessed assessment, and disk recheck in one session, then writes a self-contained verifier packet (spec.json, run.json, report.md, review.md). crucible review DIR validates the packet boundary before handoff.
  • A CI regression gate. crucible ci compares a registry's verified-latest verdicts against a sealed baseline, exits nonzero when any claim loses standing, and emits a deterministic PR-comment Markdown matrix. New in 1.2.0.
  • A refine loop that names the weakest claim. Grade each claim's measured margin, compute harmonic-mean cohesion, and re-measure across substrate rounds until the thesis is cohesively verified or the budget is spent. The loop reports the weakest axis instead of pretending a short thesis held.
  • Drift tracking across rounds. crucible drift compares the latest two witnessed assessments and classifies each claim as held, moved, improved, or regressed.
  • LLM-as-judge with the model outside the verdict. JudgeMeasure scores freeform outputs against a rubric; the judge produces a deviation once at the seam, the verdict still derives from verdict_for, and a judge that raises or returns garbage fails closed. New in 1.2.0.
  • Oracle recheck packs. crucible recheck lists descriptor-bearing measurements, writes replay templates, and validates finished replay packs against the sealed rows without opening a second verdict path.
  • A content-addressed registry with tamper detection. Every claim carries a sha256 receipt; the registry re-verifies stored claims (MATCH / MISSING / CORRUPT), checks thesis seals, rejects duplicate ids with different seals, and refuses tampered theses.
  • Typed missing-evidence explanations. Every UNVERIFIABLE verdict names the exact evidence class it lacks and the concrete next action, derived from the same pure ladder as the verdict, so the explanation can never disagree with it. New in 1.2.0.
  • Batch manifests, Markdown reports, creative measurement gates, native MCP. crucible batch runs a manifest of theses into one registry, crucible report renders deterministic Markdown, crucible measurement-gate verifies Telos creative measurement packets, and crucible mcp serves 13 tools over stdio.
  • Zero third-party runtime dependencies. The core is pure standard library, Python 3.11+.

Install

pip install crucible-bench

The distribution is crucible-bench; it installs the crucible command and the crucible package (import crucible). For the examples and the newest commands, work from a clone:

git clone https://github.com/HarperZ9/crucible
cd crucible
pip install -e ".[dev]"

Quickstart

Run the bundled demo from a clone. It registers a three-claim thesis, steelmans it, assesses it, and then shows tampering being caught:

python examples/demo.py
thesis "Binary search comparison bounds": 3 claims, seal bd7404c02eb2...
  MATCH        deviation 0 within tolerance 0.5
  DRIFT        deviation 8 exceeds tolerance 0.5
  UNVERIFIABLE claim states no falsification condition
counts: MATCH 1  DRIFT 1  UNVERIFIABLE 1
assessment seal 07dafef03f5c..., verified True
after flipping a DRIFT to a MATCH, verified False  <- caught

Then run the same thesis through the full loop into a registry:

crucible run examples/thesis-binary-search.json \
  --measurements examples/measurements-binary-search.json \
  --registry .crucible-registry
ran thesis bd7404c02eb2036e: 3 claim(s)
  steelman refutations: 3
  MATCH 1  DRIFT 1  UNVERIFIABLE 1
  assessment seal: 62f928fe21052041...
  re-derived from disk: True  {'seals_ok': True, 'thesis_ok': True, 'verdicts_rederive': True}

Add --json for the machine-readable run record, --substrate examples/substrate-binary-search.json instead of --measurements to go through the table oracle, and --bundle DIR to write a cleanroom review packet. A visual verdict surface lives at examples/crucible-demo.html.

A worked example

A thesis is claims plus falsification conditions:

{
  "title": "Binary search comparison bounds",
  "disposition": "publishable",
  "claims": [
    {
      "text": "binary search over a sorted array of 1024 elements does at most 11 comparisons",
      "falsification": "a measured worst-case comparison count above 11 for n=1024"
    },
    {
      "text": "binary search is more elegant than linear search",
      "falsification": ""
    }
  ]
}

The first claim is measurable: a measurement row binds to it by content hash, records a deviation and a tolerance, and the verdict follows. The second claim states no falsification condition, so it is UNVERIFIABLE by construction, and the assessment says exactly which evidence class is missing. Run it, bundle it, and validate the packet:

crucible run thesis.json --measurements measurements.json \
  --registry .crucible-registry --bundle reports/my-run
crucible review reports/my-run

The bundle gives a verifier only the spec and the artifact, with packet-relative paths. review fails closed on missing files, extra context, a spec.json that drifted from the run record, or a report.md that no longer renders from run.json.

To gate pull requests on the registry:

crucible ci .crucible-registry --write-baseline crucible-baseline.json   # on a known-good commit
crucible ci .crucible-registry --baseline crucible-baseline.json --out crucible-ci.md

A regression is a claim moving MATCH to DRIFT, becoming UNVERIFIABLE, or dropping out of the verified-latest set. The baseline seals its own cells, so a hand-edited baseline is rejected on load. The Markdown summary is deterministic and references the assessment seal behind every cell.

Command surface

Command What it does
crucible run THESIS --measurements F | --substrate F full loop in one session; --bundle DIR writes a review packet
crucible register / assess / steelman / measure the individual loop stages
crucible review BUNDLE validate a cleanroom packet before verifier handoff
crucible refine CONFIG rounds of substrate refinement toward cohesive verification
crucible drift DIR classify claim movement between the latest two assessments
crucible ci DIR regression gate against a sealed baseline
crucible recheck DIR inspect, template, or replay oracle measurement descriptors
crucible registry list|verify|stats|search|prune registry operations, including --require-witnessed-match
crucible verdicts DIR [--verify] list or re-derive witnessed assessments
crucible report DIR / crucible batch MANIFEST Markdown reports; manifest runs into one registry
crucible export THESIS publication-gated export; fenced material is refused at the edge
crucible measurement-gate PACKET verify a Telos creative measurement packet
crucible status / doctor / demo operator envelope, readiness checks, demo pointer (--json)
crucible mcp serve the 13 crucible tools over MCP stdio

Every command works identically from a source checkout via python -m crucible.

Extending the seams

The steelman and measure stages are seams with a stable API shape. Defaults are Null (propose nothing, measure nothing, UNVERIFIABLE), and the shipped edges include:

  • TableMeasure: offline deviation against a provided substrate table, no model.
  • SubprocessSteelman / SubprocessMeasure: configured commands over bounded JSON stdin/stdout, with timeouts, no shell strings, minimal environment, and output caps.
  • TelosMeasure: consumes telos.witnessed-artifact/v1 envelopes and re-runs the named verifier rather than trusting the carried certificate.
  • GatherDigestMeasure / IndexMeasure: sealed gather digests as evidence, and index.verification/1 records replayed against graph packs.
  • JudgeMeasure: rubric-scored LLM judging behind an injectable backend, deterministic stub in tests, null by default.

All edges map into the same Measurement to verdict_for spine. The verdict step never changes.

Documentation

Status

crucible-bench 1.2.0 covers the full loop, one-command runs, cleanroom review packets, oracle replay, registry operations, creative measurement gates, and the MCP bridge, plus what landed since 1.1.0: the CI regression gate, LLM-as-judge, missing-evidence explanations, ill-posed measurement warnings, and the MATCH-provenance gate. The test suite currently collects 330 tests. Within Project Telos, crucible is the measured-judgment layer: it consumes gather evidence and index context and emits verdict packets that telos can surface and replay.

Why it matters

Everything above reduces to one property: a verdict is a pure function of a sealed record, so anyone holding the record can recompute it, and a tampered claim, measurement, or baseline is caught by re-hashing rather than by trust.

License

crucible is fair-source: open to read, run, and build on, with commercial use reserved so the project can fund its own development. See LICENSE for the exact terms.

Work with it

python -m pip install -e ".[dev]"
python -m pytest

Keep the README, package metadata, and examples aligned with current behavior before opening a PR.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

crucible_bench-1.2.0.tar.gz (140.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

crucible_bench-1.2.0-py3-none-any.whl (94.8 kB view details)

Uploaded Python 3

File details

Details for the file crucible_bench-1.2.0.tar.gz.

File metadata

  • Download URL: crucible_bench-1.2.0.tar.gz
  • Upload date:
  • Size: 140.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for crucible_bench-1.2.0.tar.gz
Algorithm Hash digest
SHA256 95d5f2848b12b6da1c78b770fd1d56fde55863b814f6593425b38f5937a1e059
MD5 ff99b4834b1a6f26b05599a86f3eaada
BLAKE2b-256 87aa09257ae3ae0fef3cce5ef27514f8417693fd9f386d429ae27b9609ed41a0

See more details on using hashes here.

Provenance

The following attestation bundles were made for crucible_bench-1.2.0.tar.gz:

Publisher: release.yml on HarperZ9/crucible

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file crucible_bench-1.2.0-py3-none-any.whl.

File metadata

  • Download URL: crucible_bench-1.2.0-py3-none-any.whl
  • Upload date:
  • Size: 94.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for crucible_bench-1.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b696eb1a1e067987c4f9ad70cd0817f9134316c4a6769747b4345133e8dbdb69
MD5 278dd0ad881e40b128c95a4ea6554b68
BLAKE2b-256 e9f074737230e90d3b91da0156f8959e03b926cfc3df3ac801ec46e95bf40fb2

See more details on using hashes here.

Provenance

The following attestation bundles were made for crucible_bench-1.2.0-py3-none-any.whl:

Publisher: release.yml on HarperZ9/crucible

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page