skill-harness
Why this exists
I wanted to know if you could tell if a skill was any good.
That sentence is the owner's, and it is the question this repository was built to answer. The longer account, and the two commands that re-derive how much machinery the question cost, are in docs/why-this-exists.md.
A skill is a file an agent loads. Search for one that makes AI writing read less like AI writing and a dozen come back. Each costs context in every conversation, whether or not it fires. Reading the file tells you how it reads. It does not tell you whether it changes an outcome.
skill-harness runs the same task with the skill and without it, and reports what the evidence supports about the difference. The most common report is "not enough to call it."
The skills it screens live in MrBinnacle/skills. Claude Code skills are the first-class subject; the subject layer is built to take other agent ecosystems.
What does this skill cost you, and which parts of it are worth that cost?
That is the ratified wording. "Is this skill good" hides two questions. A skill has a price, paid in every conversation whether or not it fires. It has a benefit that may or may not appear when it does. The price is arithmetic on text and costs nothing to report. The benefit needs a paid comparison, and the evidence for it usually stops short of a call.
The two are measured differently, refused differently, and reported in separate fields. The instrument does not collapse them into one score.
Try the free offline skill audit
skill audit runs offline. No API key, no database, no network.
pip install skill-harness
skill-harness skill audit path/to/your/SKILL.md
It reports three properties. The cost triple: what the skill costs standing, fired, and in its side docs (aux), as arithmetic on text. A set of structural checks against Anthropic's authoring spec. An evaluability preflight: what a paid run could and could not measure about this skill today.
The output below is from that command on the committed fixture
tests/fixtures/sers/declared-synthetic-positive-control/SKILL.md, run on 2026-09-06 at
v0.3.0 in a 100-column terminal. Two lines are removed: source: (a local absolute path) and
sha256:, and the Summary line's trailing pointer to docs/concepts/why-unmeasured.md is
dropped so the line renders whole. Everything else is verbatim.
OFFLINE AUDIT — no API calls, no cost
skill: declared-synthetic-positive-control
body: 9 lines / 41 words · frontmatter keys: description, name
Structure (Anthropic authoring spec)
┌────────┬───────────────────────────────────┬────────────────────────────────────────────────────┐
│ Level │ Check │ Finding │
├────────┼───────────────────────────────────┼────────────────────────────────────────────────────┤
│ PASS │ name │ name 'declared-synthetic-positive-control' meets │
│ │ │ spec │
│ INFO │ description-unparsed-block-scalar │ description uses a multi-line YAML block scalar, │
│ │ │ which this audit's minimal frontmatter parser │
│ │ │ cannot read — description content checks skipped │
│ │ │ (UNMEASURED, not passed) │
│ PASS │ body-length │ body 9 lines (budget 500) │
│ WARN │ standing-cost-unparseable │ router listing line (frontmatter name + │
│ │ │ description) is not readable as a single-line pair │
│ │ │ — standing cost UNMEASURED (no number; a silent │
│ │ │ default would understate the per-turn tax) │
└────────┴───────────────────────────────────┴────────────────────────────────────────────────────┘
Evaluability preflight — what a paid run could measure today:
Tier-1 mechanical axes: citation_presence_per_flag, compliance_proxy, hedge_index,
structure_score, verbosity (style-shaped only).
Behavior-shaped claims (correctness, tool use, outcomes): no mechanical instrument in v0.1 → the
recorded state would be UNMEASURED, not an estimate.
Judge-graded axes: require a calibrated (judge, axis) pair — none exists in a fresh install →
UNMEASURED until you run `calibrate`.
Standing cost (mechanical): UNMEASURED — frontmatter could not be parsed well enough to count the
router listing line (see standing-cost-unparseable).
Fired cost (mechanical): raw 68 tokens · calibrated 77 tokens (x1.128, measured range
1.084-1.179) -- skill body charged when the skill runs and its body is read.
Aux cost (mechanical): raw 0 tokens · calibrated 0 tokens (x1.128, measured range 1.084-1.179) --
other documentation files beside the skill (progressive disclosure); 0 when the skill directory has
none.
Summary: 2 pass · 1 warn — UNMEASURED is a recorded state, not a failure.
Clause evidence: UNMEASURED (no_extraction: clause evidence requires an extraction output; produce
one with skill init --out <file>)
The parser could not read the description, so the audit says so and skips the check. It does
not pass a field it never read. The standing cost has no number for the same reason, and the
audit prints UNMEASURED rather than a default.
--strict exits 1 on warnings, for CI. From v0.3.0 the CLI sets UTF-8 on its own output
streams; on v0.2.3 a Windows console needs PYTHONUTF8=1 set first or the command exits 1 with
UnicodeEncodeError.
What it has found so far
Zero production-skill KEEPs. The full keep lane has fired end to end once, on 2026-07-27. The subject was a declared synthetic positive control: a skill written to carry an invented fact, so the effect exists by construction. It returned KEEP at 8/8 with the skill against 0/8 without, posterior probability of a win 0.99 (SERS receipt, 2026-07-27). That run shows the instrument fires when an effect is present. It says nothing about whether any real skill is worth its slot.
The most common result is that the model already does the task without a skill. On two deliberately hardened tasks a frontier agent passed 14 of 14 no-skill runs. Nothing was left for a skill to improve, so nothing could be measured. That is a finding about the task, and it is written up in the double-ceiling case study.
The paired run in that study, July 2026, cost about $6.17 and returned the pre-registered
NO-GO: an apparatus check, not a measurement of benefit. The receipt records it as one
(double-ceiling-nogo-2026-07-09.json).
None of that is a scheduling accident. A sized benefit run launches only on the first task whose no-skill screen returns a pass rate below 1. Every production skill screened so far ceilings at 1: the model passes every attempt without it (observation ledger).
Why it refuses
Comparing a skill against nothing is noisier than it looks. On a 60-trial arc of identical agentic coding tasks, run-to-run output-token variation measured CV ≈ 17.6% (RMS across cells; mean-of-cells 14.6%, median 10.4%) on an Opus-class model (findings record). At that coefficient of variation, a three-runs-a-side comparison cannot separate differences under roughly 30–40% from noise, and three runs a side is what most published skill comparisons use. Hand-picked tasks tilt the result before anything runs. Pass/fail test banks price what a skill costs and skip what it does. The findings record carries an evidence grade on each claim.
The design rule follows from that. A figure that is not there is stated as a typed refusal, never filled in. No placeholder zero, no free-typed excuse, no estimate standing in for a measurement.
Three consequences follow from the rule.
One. Every paid comparison has a control arm. With and without, never a score in a vacuum, because a score in a vacuum cannot show that the model did not need the skill.
Two. When the evidence cannot carry a call, the answer is UNMEASURED with a reason from a
fixed list of eight: no_data, inadmissible, underpowered, falsifying_case_missing,
budget_exhausted, falsifying_case_stale, fdr_correction_failed, mechanical_vacuous
(src/skill_harness/aggregation/status.py). Which kind of not-knowing is more information
than not-knowing alone. Definitions:
docs/concepts/why-unmeasured.md.
Three. Evidence passes a gate before it enters an aggregate, and the gate result is snapshotted at write time in an append-only store. Data that fails the evidence-admissibility gate is kept and never counted. A judge-graded result counts only where that judge has been calibrated on that axis first. Calibration swaps answer order to cancel position bias, controls for length, defends against injection, and measures agreement with a human.
The same rule points inward. Two of the instrument's own weak points are measured, and both numbers stay on the front page:
Extraction repeat-variance: MEASURED for one skill — three repeat extractions of the same
SKILL.md returned 29/33/34 clauses, so clause counts are not stable run to run and
nothing downstream is allowed to key on clause position
(#152).
Vacuity-flag precision: MEASURED at 0.972 by blind cross-family adjudication over 106
adjudicated rows, and that figure is flag-level only
(#153). When the adjudicators also
had to agree on which kind of vacuity, kind-precision 0.835: not_a_directive matched 77/77,
while weak_directive matched 4/20. The vacuity-flag detector's recall is UNMEASURED: the
unflagged clauses were never adjudicated.
What it measures, and what it refuses to
The answer comes back as one of three verdicts: KEEP, CUT, or CAN'T-TELL-YET.
Which of those a skill is eligible for depends on its registered value class, not on the
numbers alone. A CUT says why: subsumed (the model was already doing it), no_lift (the
model needed help and the skill did not deliver it), or harmful.
The value-class guard sits on that. A skill can exist to stop one specific wrong move. A
model that passes without such a skill has not shown the skill is useless; it has shown the
trap did not come up. So subsumed is a CUT only for skills registered as
TRANSFORMATIVE_LIFT, the class whose whole claim is lift above the bar. Every other class
reclassifies to CAN'T-TELL-YET, because this is the wrong instrument for that kind of skill,
not a verdict on it.
Two skills from the collection moved that way when the guard landed: append-only-evidence-design
(calibration) and a hardened git-pull-rebase-trap (trap-discipline). Under the pre-guard rule
both returned CUT (subsumed), each at a no-skill pass rate of 1.00; the value-class guard
reclassified both to CAN'T-TELL-YET
(receipts). The
pre-guard CUTs stay in the record as dated output, not edited into agreement.
The other half has never fired: a paired run sizing how much a skill helps once the model is known to need help. By design, a sized benefit run launches only when a screen returns a sub-1 pass rate, and none has.
Measuring for real
skill-harness skill init path/to/SKILL.md --execute # extract testable claims
skill-harness run ablation <skill_id> --execute # the with/without comparison
skill-harness run evaluate-skill <skill_id> # aggregate to a verdict
ANTHROPIC_API_KEY or OPENROUTER_API_KEY. Every run subcommand is dry-run by default;
--execute is required to spend, and a per-run cap and a daily cap sit on top.
skill init is the exception: clause extraction is a model call in both modes, and
--execute decides only whether the result is persisted to the evidence DB. Without a key
it exits 1 before any call.
Reproduction scripts:
examples/.
The reporting vocabulary is a published standard
The Skill Efficacy Reporting Standard (SERS) fixes the vocabulary: verdicts, refusal
reasons, the cost triple, the evidence-admissibility statuses, and the instrument identity.
Instrument identity is the model pin and prompt fingerprint that stamp which generation
produced a figure. SERS is a JSON Schema plus a prose companion, in
docs/sers/.
SERS is separate from this tool's internals on purpose. Another harness can emit conforming reports without adopting anything here. CI checks that this repository's own receipts validate against the schema, that the schema's enums match the code's, and that deliberately poisoned receipts are rejected.
Models change underneath every figure, so every figure has a shelf life. Instrument identity is a required field: two numbers from two generations are visibly non-comparable rather than averaged.
What this isn't
It is not the most featureful skill benchmarker available. If you want the most featureful skill benchmarking today, adewale's skill-eval-harness is the closest neighbour and is further along on more axes. A few of its disciplines are on the adoption list, with attribution. For comparing prompts and configurations rather than skills, promptfoo is the mature choice. For evaluating models and agents, Inspect is the institutional one.
Reach for this one when the question is whether the number deserves to exist at all.
This repository makes no first-mover claim. The positioning was checked against primary sources before it was written, and the first-or-only claims failed that check (#39). Two claims carry the most weight: the pre-spend eligibility gate, and the rule that thresholds are ratified from enumerated tables rather than authored by hand. Both carry claim-status labels tied to an external review plan (#45). They change by dated amendment, never silently.
The other half
The verdicts land in a second repository: MrBinnacle/skills, a small collection. There, each skill carries its own dated evidence record and controlled results are read from that skill's record, not a front-page roll-up. Skills are re-screened when a major model ships and publicly retired, with the record intact, once the model no longer needs them or a platform change meets a pre-registered trigger. Each retirement is made against its stated criterion.
The two repositories run on one rule, pointed at two different targets. This one does not state a number the evidence does not support. That one does not keep a skill the evidence no longer supports.
Dig deeper
- The receipts, rendered — the SERS receipts as a browsable site, one page per screened skill, cost triple beside the evidence grade. It renders the SERS instances only; the Markdown index below is the citable surface for every kind.
- Measurement receipts index
— every case study, finding, observation, assurance report, ratification, SERS instance, and
the
skill audit --extractionjoin surface: what each claims and what each refuses to claim. - Why this exists — how a non-specialist ends up building a measurement instrument, and the loop that made it possible.
- The double-ceiling case study — the run where there was nothing left to measure.
- The ablation that caught its own author — three pre-spend catches before a contaminated result could ship.
- When ablation measures the wrong layer — if a discipline fires in a hook, ablating the skill text says nothing about the discipline.
docs/PRD.md— the full specification: evidence model, oracle tiers, gate rules, CLI surface.- The observation ledger — per-record screen history, annotated rather than rewritten.
Status: v0.3.0 on PyPI. Not every older screen record is yet in the evidence store; the observation ledger shows the evidence behind each record.
MIT licensed. Issues and PRs welcome:
CONTRIBUTING.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file skill_harness-0.3.0.tar.gz.
File metadata
- Download URL: skill_harness-0.3.0.tar.gz
- Upload date:
- Size: 1.5 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0897cf94184ba027871c0671f6cc1b1d61a156941db6cb0aa8684525374de1c2
|
|
| MD5 |
743342eee6f9d50a610c96a1f9e8e3a8
|
|
| BLAKE2b-256 |
a99c6bbe63cbf05b38ec7cf3ca2698e8e0ba8f4a2cec350456bcf7d591b8a674
|
Provenance
The following attestation bundles were made for skill_harness-0.3.0.tar.gz:
Publisher:
publish.yml on MrBinnacle/skill-harness
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
skill_harness-0.3.0.tar.gz -
Subject digest:
0897cf94184ba027871c0671f6cc1b1d61a156941db6cb0aa8684525374de1c2 - Sigstore transparency entry: 2744927960
- Sigstore integration time:
-
Permalink:
MrBinnacle/skill-harness@140b0e379246ce71421d418ca736d0f1463225b9 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/MrBinnacle
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@140b0e379246ce71421d418ca736d0f1463225b9 -
Trigger Event:
release
-
Statement type:
File details
Details for the file skill_harness-0.3.0-py3-none-any.whl.
File metadata
- Download URL: skill_harness-0.3.0-py3-none-any.whl
- Upload date:
- Size: 502.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5575cee3e7da4d616dddf00c8750699528ef5f41d41c8c2be38b098509988a3b
|
|
| MD5 |
5e1f35385727050a733400eeaf3ea52d
|
|
| BLAKE2b-256 |
3a0a471eec6c1f0fe8b770a4314bdec9e878da444cf18e79dfcb35a5226a19ce
|
Provenance
The following attestation bundles were made for skill_harness-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on MrBinnacle/skill-harness
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
skill_harness-0.3.0-py3-none-any.whl -
Subject digest:
5575cee3e7da4d616dddf00c8750699528ef5f41d41c8c2be38b098509988a3b - Sigstore transparency entry: 2744928667
- Sigstore integration time:
-
Permalink:
MrBinnacle/skill-harness@140b0e379246ce71421d418ca736d0f1463225b9 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/MrBinnacle
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@140b0e379246ce71421d418ca736d0f1463225b9 -
Trigger Event:
release
-
Statement type: