Skip to main content

Eval Conductor

CI

The integrity layer for your evals. Your evals already produce numbers — from claude plugin eval, pytest, promptfoo, or your own scripts. This proves them: that the verdict is real, the bar was fixed before the run, and the decision is tamper-evident. It sits next to MLflow, LangSmith or Braintrust; it doesn't replace them. It answers the one question none of them do:

Can you prove this eval result — or do we just trust it?

Open source, MIT. stdlib + PyYAML only. No telemetry, no accounts, no infrastructure. The only outbound call is optional — /conductor --verify-evidence checks one of your own CI runs against the GitHub API — and receipts are signed with a key you generate locally. Dogfooded on its own 178 tests.

The thesis: trust an AI verdict only by what it can prove.

Start in 60 seconds — no config

/eval-prove

Runs the tests you already have, emits a signed, tamper-evident receipt that they passed at this commit, and prints a badge to paste in your PR:

PROVEN: 142/142 tests pass (100%)   receipt a1b2c3…
![evals](https://img.shields.io/badge/evals-proven%20142%2F142-brightgreen)

The first run captures your current pass rate and test count as the bar. A later run that quietly drops the pass rate — or deletes tests to go green — gets BLOCKED. No registry, no setup. When you outgrow it, /eval-init installs the full gate.

The three things no other eval tool does

A green eval is ephemeral: nothing proves what passed, under which bar, or that a failure was ever produced. Eval Conductor closes that with three controls that compose into one claim the incumbents can't make.

1. The bar can't silently weaken. A registry diff-lint compares your gate against its own git history. Lower a threshold, flip a direction, downgrade a blocking tier, widen a band, or delete a metric, and CI fails unless you committed a written reason next to the change. "Never edit a threshold to turn a red gate green" — made mechanical.

2. The bar was fixed before the result. /eval-seal hashes your registry (and band floors) into a signed, hash-chained seal — containing no scores — before you run. At gate time the live bar is re-hashed against it; a bar that moved after sealing is caught. This closes p-hacking your own gate: choosing the threshold after you've seen the number. (The diff-lint can't catch that on a brand-new gate — there's no history to diff. The seal is that history, on purpose.)

3. The decision is tamper-evident proof. Every PROMOTE/BLOCK emits an HMAC-signed, hash-chained receipt binding the decision to the exact registry, candidate and baseline bytes it was computed from — stamped with whether it honored the pre-registered bar. Alter or reorder a receipt, or remove one from the middle, and the chain breaks. Removing the latest receipt leaves a shorter valid chain — that case is caught not by the hash link but by distribution: commit and push your receipts and every clone (and your git history) holds the fuller chain, so a truncation is visible. "It passed" becomes cryptographic proof it passed, unaltered, under this bar, at this time.

We proved this on ourselves: we p-hacked our own gate on purpose, and the receipt caught it.

Fail-closed by design

Every rule exists because a real gate was seen passing a bad run without it. Pointed at itself during its own build, this gate caught its author closing steps on placeholder evidence three times (preserved in the committed ledger) and found that NaN scores silently passed every threshold check.

  • NaN / non-numeric score → BLOCK, even on monitor-only metrics. A broken instrument is not a passing run.
  • Missing metric → BLOCK. Deleting a metric is never the easy way to pass.
  • Stale baseline / tampered registry → BLOCK. The registry hash is recomputed from disk, so editing it after the baseline can't ride a stale attestation.
  • Partial eval run → refused. A run cut short by a cost ceiling gates nothing.
  • Zero tests executed → 0.0, not 1.0 in the gate; and /eval-prove refuses it outright (CANNOT PROVE) — a run that measured nothing is never a green badge.
$ /eval-gate

metric          block      cand     base    band  status  reason
test_pass_rate  hard     0.6667      -    0.0000  BLOCK   threshold breach: 0.6667 vs 1.0 (higher_better)
test_count      hard     3.0000      -    0.0000  PASS    ok
test_failures   hard     1.0000      -    0.0000  BLOCK   threshold breach: 1.0000 vs 0.0 (lower_better)

DECISION: BLOCK            (exit 1 — CI stops here)

Works with what you already run

You produce Adapter Gated on
pytest --junitxml (or any JUnit: jest, go test…) junit pass rate, test count, failures
claude plugin eval --json plugin-eval score, with-vs-without delta, min case, cases passed
promptfoo eval -o out.json promptfoo pass rate, total, failures
DeepEval, Ragas, Inspect, your own script scores any flat {metric: number}

The gate judges the numbers it's given. Measuring the right thing is still your job — the eval-gates skill is about how.

Commands

Command What it does
/eval-prove Start here. Zero config: prove your existing tests with a signed receipt + a pasteable PR badge. First run sets the bar; a later regression blocks.
/eval-seal Pre-register the bar: seal the registry + band floors before a run, so a decision can prove it was judged against a bar fixed in advance.
/eval-verify Walk the receipt chain offline: structure, every signature, every decision vs. its own verdicts. Any tamper, fork, or mismatch → exit 1 naming the receipt.
/eval-init Install the full gate into a repo: registry, gate CLI, adapters, baseline runner, BLOCK-fixture guide, CI workflow with the producer pre-filled.
/eval-gate Candidate vs. baseline under the registry → per-metric verdict table → PROMOTE (exit 0) / BLOCK (exit 1).
/eval-baseline Run your eval N times (default 10) and measure noise bands per metric — mean ± 2σ, or a robust median ± MAD for band_method: mad. Warns on low N or a zero band.
/eval-explain Explain a decision in plain English: which metric, breach vs. regression beyond noise, the one legitimate fix — and the theater moves named.
/eval-import Convert scores you already have into candidate.json: junit, plugin-eval, promptfoo, scores.
/conductor The 27-step evidence-gated build ledger: steps close only on a real, green CI run (GitHub API–verified); placeholder evidence refused; reversals logged, not erased.

What this proves — exactly, no more

Honesty is the product, so here are the limits in plain sight:

  • A receipt proves a decision is unaltered since signing under a key your team controls. It does not identify who signed, and it does not prove the candidate scores were produced honestly (that's the gate's and the adapters' job). Coverage is only as good as the chain is distributed — commit and push your receipts; a local-only chain is a claim, not evidence.
  • A seal proves the bar is byte-identical to what was sealed and — via pushed git history — committed before the result. It does not prove the bar is strict enough. A pre-committed weak bar is still weak. A shared-secret key can't stop its holder backdating the timestamp, which is why ordering rests on pushed history, not the clock.
  • A green /eval-prove badge proves "not worse than before," not "good." The badge always shows the test count so it can't imply coverage it doesn't have.

Install

Requires Claude Code, Python 3.10+, and PyYAML (pip install pyyaml). No telemetry and no accounts; the only outbound call is the optional /conductor --verify-evidence check against the GitHub API.

/plugin marketplace add Rmasood1122/eval-conductor
/plugin install eval-conductor@eval-conductor-marketplace

Or from your shell: claude plugin marketplace add Rmasood1122/eval-conductor && claude plugin install eval-conductor@eval-conductor-marketplace. Confirm with /plugin list; if /eval-prove isn't recognized, start a fresh session — plugins load at startup.

Skills

  • eval-gates — designing a registry, picking blocking tiers, writing per-hard-metric BLOCK fixtures, and the failure modes naive gates miss.
  • evidence-discipline — the conductor method: evidence or it didn't happen, honest reversals, fake-evidence patterns to watch for in your own work.

Provenance

The tamper-evident receipt engine (canonicalization + fork-detecting hash chain) is fused from titan-gate, signed receipts for AI-assisted commits, by the same author. Same thesis, different layer.

License

MIT.

Metadata

Release files for eval-conductor 1.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for eval-conductor 1.4.1
File Size Uploaded
eval_conductor-1.4.1.tar.gz 84.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for eval-conductor 1.4.1
File Interpreter ABI Platform
eval_conductor-1.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 150.1 kB

Release files / eval_conductor-1.4.1.tar.gz

Download URL eval_conductor-1.4.1.tar.gz
Size 84.5 kB
Tags Source
SHA-256 checksum
How to use checksums
1fc7b4fdfe329e42642c3016ee0526e2569c435a7ad3723afbdbcee684284ef8
BLAKE2b-256 checksum
How to use checksums
e95ed725a0b852c14bd3b72dface4bf887369419a56f368712465022cb5ed3c2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / eval_conductor-1.4.1-py3-none-any.whl

Download URL eval_conductor-1.4.1-py3-none-any.whl
Size 65.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8c1f43d524d4019a33ded6defa0745a65219b37f755562c69f20b3e7b9e708d6
BLAKE2b-256 checksum
How to use checksums
c56e00f46fd8794e2625ba3d385e83139f95dd7d149cea980b284e74e03b9b0b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

1.4.1 This release

2 release files

1.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page