Skip to main content

An open spec for A/B benchmarking skills via declarative test suites.

Project description

skillevaluation

Does your skill actually make the agent better? Prove it — with measured before/after numbers.

CI PyPI Python License

A skill is just a folder — a SKILL.md plus some attachments. It's easy to write one and assume it helps. skillevaluation lets you measure the help: write a small eval.yaml next to your skill, and a runner executes each test case twice — once with the skill loaded, once without — then hands you a clear A/B delta on pass rate, speed, tokens, turns, and tool calls.

No more "I think this skill is good." Now you can say "this skill lifts pass rate 40 points and cuts tokens 43%" — and back it with reproducible cases.


The payoff

Here's the bundled gdpr-pii-classifier example — the same five cases, run with the skill and without it:

Dimension Without skill With skill Delta
Pass rate 40% 80% +40 pts
Avg tokens 3,210 1,840 −43%
Avg turns 8.2 4.6 −44%
Avg duration 22.8s 14.2s −38%
Avg tool calls 5.4 3.0 −44%

The skill more than doubled the pass rate and made the agent faster and cheaper. That's exactly the kind of claim skillevaluation is built to produce.

Numbers above are illustrative of the example's shape — your real deltas depend on your agent runtime and model.


How it works

  1. Write eval.yaml next to your SKILL.md — a handful of declarative cases (a prompt, plain-English expectations, and optional shell validators).
  2. A runner executes each case twice — once with the skill loaded (the with arm), once without (the without arm).
  3. You get measured deltas — each case is classified (flip_to_pass, pass_kept, …) and aggregated into per-dimension lift.
                    ┌─ with skill ────▶ pass? + metrics ─┐
   each case ──────▶┤                                    ├──▶ outcome ──▶ aggregate deltas
                    └─ without skill ─▶ pass? + metrics ─┘

Quickstart

pip install "skillevaluation[runner]"

1. Describe what "better" means. Drop an eval.yaml beside your SKILL.md:

# eval.yaml
cases:
  - name: tracks_with_id
    prompt: "Classify these schema fields and write JSON to output.json: email, ip_address, name, age."
    expectations:
      - "The response classifies email as PII"
      - "The response identifies ip_address as pseudonymous (not PII)"
    validators:
      - cmd: "jq -e '.email.category == \"PII\"' output.json"
        label: "email categorized as PII"

See the full five-case suite in examples/gdpr-pii-classifier/eval.yaml.

2. Run the A/B benchmark. One command, your own API key — nothing leaves your machine:

export ANTHROPIC_API_KEY=sk-ant-...   # or OPENAI_API_KEY / GEMINI_API_KEY
skillevaluation run ./skills/gdpr-pii-classifier --model claude-haiku-4-5

Each case executes twice — once with the skill loaded, once without — then both arms are graded with the same assertions: expectations via an LLM judge, validators in a sandboxed workspace. You get the delta table on stdout plus a results.json that validates against the packaged wire schema. The without-skill arm is cached locally (it doesn't depend on the skill), so re-runs while you iterate on SKILL.md cost half.

3. Gate it in CI.

- run: skillevaluation run ./skills/gdpr-pii-classifier --fail-on-verdict fail --min-delta-pts 10

Useful flags: --adapter mock (free, networkless plumbing dry-run) · --adapter claude-code (experimental: drives your installed Claude Code, so turn/tool-call deltas come from a real agent loop) · --judge-model (use a cheap judge, e.g. gemini-3.5-flash) · --trajectories DIR (write per-arm canonical transcripts) · --json · --export-url (POST the results document to any collector).

Honest-metrics note: the default llm adapter is a single-shot completion — pass-rate, token, and duration deltas are real; turns/tool_calls are trivially 1/0. Use an agent-runtime adapter (or implement AgentAdapter for your own stack) when those dimensions matter.

Scoring as a library. Bringing your own harness? The same delta math is importable directly:

from skillevaluation.outcomes import classify_outcome
from skillevaluation.aggregation import CaseResult, CaseMetrics, compute_run_aggregates

results = [
    CaseResult(
        case_name="tracks_with_id",
        outcome=classify_outcome(with_passed=True, without_passed=False),
        with_skill=CaseMetrics(passed=True,  duration_ms=14200, turns=4, total_tokens=1840, tool_call_count=3),
        without_skill=CaseMetrics(passed=False, duration_ms=22800, turns=8, total_tokens=3210, tool_call_count=5),
    ),
    # ... one CaseResult per case
]

agg = compute_run_aggregates(results)
print(agg.pass_rate)   # {'with_skill': 1.0, 'without_skill': 0.0, 'delta_pts': 100.0}
print(agg.to_dict())   # full per-dimension JSON, matching the wire schema

What actually runs the agent? The bundled reference runner — skillevaluation run executes every case A/B through a pluggable adapter. Building a runner in another language (or a hosted one)? It's an open spec: start at the runner contract and verify against compatibility-tests/. Hosted execution, run history, and rankings are platform concerns — DecimalAI runs this same spec server-side.

Status: v0.2.0, pre-1.0. The format is stable enough to build on, but APIs may shift before v1 — changes are logged in CHANGELOG.md.


What's in the box

A typed Python reference implementation. The core is dependency-light (only needs PyYAML); the runner's HTTP pieces live behind the [runner] extra:

Module What it does
skillevaluation.parser Parse + strictly validate eval.yaml
skillevaluation.outcomes Classify each case: flip_to_pass / pass_kept / fail_kept / flip_to_fail / error
skillevaluation.aggregation Per-dimension delta math, with an honest apples-to-oranges skip rule
skillevaluation.baseline Baseline-cache key derivation (skip re-running an unchanged without arm)
skillevaluation.trajectory.format_v1 Canonical agent-session rendering, so different runners' LLM judges agree
skillevaluation.resources The packaged spec/ + schemas/ — validate results offline, no GitHub fetch
skillevaluation.runner The reference runner: A/B orchestrator, reference LLM judge + structural fast path, sandboxed validators, local baseline cache
skillevaluation.runner.adapters The invocation seam: direct-LLM (supported), mock (deterministic), Claude Code (experimental) — or implement AgentAdapter for your runtime
skillevaluation CLI run (delta table + results.json + CI gates) and validate

Use it as a spec, not just a library

skillevaluation is an open spec, so any tool — in any language — can produce interoperable results. If you're building your own runner, start here:

Deliberately out of scope: live traffic-split experiments, external eval-score webhooks (DeepEval/LangSmith), catalog ranking or publish-gate policy, and the exact LLM-judge prompt wording (the contract is specified; the prompt is your choice).


Contributing

Contributions are genuinely welcome — especially new conformance cases that catch an edge the golden suite misses. See CONTRIBUTING.md. Dev setup is the usual:

git clone https://github.com/decimal-labs/skillevaluation
cd skillevaluation
pip install -e ".[dev]"
pytest

License

Apache 2.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

skillevaluation-0.2.1.tar.gz (103.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

skillevaluation-0.2.1-py3-none-any.whl (78.9 kB view details)

Uploaded Python 3

File details

Details for the file skillevaluation-0.2.1.tar.gz.

File metadata

  • Download URL: skillevaluation-0.2.1.tar.gz
  • Upload date:
  • Size: 103.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for skillevaluation-0.2.1.tar.gz
Algorithm Hash digest
SHA256 212844f599f968aa8aa533b293856d99243d2fa1c32f572b4cd8dbadcfb1a1ce
MD5 3214a0518c1e68e226bdd5ce929736d7
BLAKE2b-256 e44a6e0c36a140092a2efac208d6b4d27f4eb62b15123384de267b57564c60cc

See more details on using hashes here.

File details

Details for the file skillevaluation-0.2.1-py3-none-any.whl.

File metadata

File hashes

Hashes for skillevaluation-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 a44477712b34be91bde8e5ad5f9af2e87883b825622ff6f35e459f227ddb9d99
MD5 714de1bb5606590d37fefc4da9373824
BLAKE2b-256 08fcd048e94e44881f063b5d2017ef34d4c6d63457adca0bde87ad9f2f7da18b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page