CLI for evaluating Claude Code skills and AI agents
Project description
Caliper
pytest for agent skills. Write a task in YAML, run it k times, get a reliability score.
npx skills@latest add edonadei/caliper
Agent skills are hard to test. A skill that works on your machine, on this prompt, today, might fail tomorrow after a model update or a one-line prompt edit. Caliper makes reliability measurable: define what success looks like, run the skill repeatedly, and get a pass@k score you can track over time.
Use Caliper to answer questions like:
- Did my prompt edit actually improve the skill?
- Is the skill doing the work, or would the base agent pass without it?
- Does it still pass the workflows it passed last week?
- Which backend — Claude Code or Codex — runs this skill more reliably?
How it works
.eval.yaml spec
│
▼
Harness ──── runs your skill against the agent (Claude Code / Codex / API)
│
▼
Judge ──── LLM autorater and/or deterministic Python assertions
│
▼
pass@k score + saved transcript
Each attempt runs in an isolated temporary home with no session history. Results are saved as JSON you can inspect and diff later.
Quick start
1. Install
Install the evaluate-skill agent skill — it handles everything from inside your agent:
npx skills@latest add edonadei/caliper
Or install the CLI directly if you prefer running evals from the terminal:
pipx install caliper-eval
Requires Python 3.10+.
2. Create a spec
# my-skill.eval.yaml
skill:
path: ./SKILL.md
backend: claude-code
judge:
backend: claude-code
tasks:
- name: Writes a greeting file
cleanup: rm -f /tmp/hello.txt
prompt: "Write 'hello world' to /tmp/hello.txt"
expect: "A file at /tmp/hello.txt containing 'hello world' was created."
assert: |
from pathlib import Path
assert Path("/tmp/hello.txt").read_text().strip() == "hello world"
3. Run it
If you installed via the skill, ask your agent:
/evaluate-skill run my-skill.eval.yaml --k 3 --baseline
Or from the terminal if you installed the CLI:
caliper run my-skill.eval.yaml --k 3 --baseline
4. Read the output
CALIPER - my-skill - k=3 - claude-code
ID Task k (3) pass@k
task-1 Writes a greeting file 3/3 100% PASS
With skill 100% ####################
No skill 70% ##############------
Delta +30% up
Results saved to .caliper/results/my-skill/2026-06-19T14-23-01Z.json
Browse past results anytime:
/evaluate-skill list
/evaluate-skill report my-skill
Core concepts
| Term | What it is |
|---|---|
| Spec | A .eval.yaml file that describes the skill, judge, and tasks to run |
| Backend | The agent that executes the skill (claude-code, codex, claude-api, openai-api) |
| Judge | What decides pass/fail — an LLM reading the transcript (expect:), Python assertions (assert:), or both |
| pass@k | Reliability score: run k times, measure how often the skill succeeds |
| Baseline | Re-run the same tasks without the skill to prove the skill is doing the work |
| Attempt | One isolated run of a single task — fresh temporary home, no session history |
Choosing a backend
| Backend | Requires | Best for |
|---|---|---|
claude-code |
Claude Code CLI installed and authenticated | Testing Claude Code slash-command skills |
codex |
Codex CLI installed (npm install -g @openai/codex) |
Testing Codex skills |
claude-api |
ANTHROPIC_API_KEY env var |
API-backed agents, no CLI needed |
openai-api |
OPENAI_API_KEY env var |
OpenAI API agents |
The agent backend and judge backend are independent — you can test a Codex skill with a Claude judge, or any other combination.
Claude Code setup
Install and authenticate the claude CLI. backend: claude-code uses your existing Claude Code auth — no extra configuration needed.
For backend: claude-api:
export ANTHROPIC_API_KEY=...
Codex setup
npm install -g @openai/codex
codex login
backend: codex calls codex exec. It does not fall back to the OpenAI API. If the Codex desktop app is installed, Caliper prefers the app-bundled binary over codex on PATH. Set CODEX_CLI_PATH to force a specific binary.
For backend: openai-api:
export OPENAI_API_KEY=...
Check installed CLI versions:
caliper update-cli --check
Recommended workflow
- Create a spec for one behavior you care about.
- Run with
--k 1while iterating on the spec. - Add
assert:for facts an LLM judge might guess wrong (files, JSON, command output). - Move to
--k 3or higher once the task is stable. - Add
--baselineto prove the skill is making a difference. - Commit the spec alongside the skill so contributors can run the same eval.
caliper run my-skill.eval.yaml --k 3 --baseline --verbose
Spec format
skill:
path: ./SKILL.md # path to the skill file (optional for baseline-only runs)
backend: claude-code # claude-code | codex | claude-api | openai-api
model: <model-name> # optional model override
judge:
backend: claude-code # claude-code | codex | claude-api | openai-api
model: <model-name> # optional model override
sandbox:
extra_path:
- ./bin # prepended to PATH inside each attempt
forbidden_files:
- ".*\\.eval\\.yaml$" # prevents agent from reading the spec
- "./.caliper/.*" # prevents agent from reading saved results
tasks:
- name: Short task name
setup: <shell command> # optional, runs before each attempt
cleanup: <shell command> # optional, always runs after each attempt
prompt: <prompt sent to the agent>
expect: <natural-language success condition>
assert: |
# optional inline Python assertion
assert True
- name: Task with external assertion script
prompt: "Generate a report"
assert: ./assertions/check_report.py
Each task needs at least one of expect or assert. Task IDs are assigned automatically as task-001, task-002, and so on.
Judging
LLM autorater (expect:)
The judge backend reads the full attempt transcript and decides whether the expect condition was met. When the backend captures tool-call traces (Claude Code, Codex), those traces are included — the judge can verify things like "the agent used tool X" without relying on the final text alone.
judge:
backend: claude-code
Deterministic assertions (assert:)
Python assertions run locally. Use these for facts the LLM judge might guess:
- file exists / exact file contents
- JSON / schema validity
- command output
- images or screenshots
- repository state
tasks:
- name: Writes an output file
cleanup: rm -f /tmp/out.txt
prompt: "Write hello world to /tmp/out.txt"
assert: |
from pathlib import Path
path = Path("/tmp/out.txt")
assert path.exists(), "Output file was not created"
assert path.read_text().strip() == "hello world"
When both expect and assert are present, both must pass.
--judge flag
| Flag | Behavior |
|---|---|
--judge autorater (default) |
LLM judge evaluates expect |
--judge script |
Runs assert: always; also runs LLM judge if expect is present |
CLI reference
| Command | Description |
|---|---|
caliper run <spec> |
Run an evaluation spec |
caliper new [name] |
Create a new spec with the interactive wizard |
caliper validate <spec> |
Validate a spec file |
caliper list [spec] |
List specs and saved runs |
caliper report <spec-or-result> |
Re-render saved results |
caliper install-skill <backend> |
Install the bundled evaluate-skill into Claude Code or Codex |
caliper update-cli [backend] |
Check or update installed agent CLI versions |
caliper run flags
| Flag | Default | Description |
|---|---|---|
--k INT |
3 |
Attempts per task |
--baseline |
off | Also run each task without the skill |
--judge MODE |
autorater |
autorater or script |
--workers INT |
4 |
Parallel task workers |
--timeout INT |
120 |
Seconds per attempt |
--model MODEL |
— | Override skill.model |
--verbose |
off | Show per-attempt judge reasoning |
--output PATH |
— | Also save results JSON to a specific path |
Scoring
For each task:
pass@k = 1 - (1 - successes / k) ^ k
The aggregate score is the average task pass@k. With --baseline, Caliper runs the same tasks without the skill and reports the delta.
Agent skills
The repo ships two agent skills. Install both with:
npx skills@latest add edonadei/caliper
evaluate-skill — run and manage evals
Create, validate, run, and summarize evals from inside your normal workflow — no separate terminal needed. The skill installs Caliper automatically if it's missing.
Or, if you already have Caliper installed and want to wire up the skill manually:
caliper install-skill claude-code
caliper install-skill codex
Preview without writing files:
caliper install-skill claude-code --dry-run
Then use it in Claude Code:
/evaluate-skill run my-skill.eval.yaml --k 3
/evaluate-skill validate my-skill.eval.yaml
Or in Codex:
Use the evaluate-skill skill to run my-skill.eval.yaml with k=3 and summarize the result.
grill-skill — create evals interactively
Don't have evals yet? grill-skill guides you through creating them. It reads your SKILL.md, interviews you about what good behavior looks like, and generates a 3-task spec (happy path, edge case, adversarial). Then it runs the eval and loops — k=1 to validate, k=3 to measure, baseline before you commit.
/grill-skill ./my-skill/SKILL.md
No path needed if you're already in the skill's directory:
/grill-skill
If an .eval.yaml already exists next to your skill, grill-skill reads the existing tasks and interviews you about gaps instead of starting from scratch.
Project layout
caliper/
commands/ CLI command implementations
harness/ Agent execution backends (Claude Code, Codex, API)
judge/ LLM and script judging implementations
schema/ Eval spec and result models
runner.py Evaluation orchestration
skills/
evaluate-skill/ Agent skill for running Caliper from Claude Code or Codex
grill-skill/ Agent skill for creating and iterating on evals interactively
tests/ Pytest coverage for harnesses, judges, and runner behavior
Contributing
Good first areas:
- add example evals for real skills
- improve backend error messages
- add deterministic assertion helpers
- expand tests for harness and judge behavior
- improve result reporting and summaries
- document common setup problems for Claude Code and Codex
Before opening a pull request:
pip install -e ".[dev,openai]"
pytest
ruff check .
caliper validate skills/evaluate-skill/evaluate-skill.eval.yaml
When changing behavior, include a test or an eval fixture that demonstrates the expected outcome. Keep backend-specific logic isolated to the relevant module under caliper/harness/ or caliper/judge/.
Troubleshooting
codex judge failed: model ... is not supported
The model name is not available to your Codex account. Use a model that codex exec --model <name> accepts.
codex CLI not found
npm install -g @openai/codex
claude command not found
Install and authenticate Claude Code, or switch the backend to codex, claude-api, or openai-api.
A task passes only because of assert:
When a task has only assert:, no LLM judge runs. Add expect: if you also want an LLM to evaluate the transcript.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file caliper_eval-0.2.2.tar.gz.
File metadata
- Download URL: caliper_eval-0.2.2.tar.gz
- Upload date:
- Size: 340.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9c74563f8d7daf27ee5d5271f166d852c3a1e19d6aa1c9e16be235d0ff74d191
|
|
| MD5 |
ad73681525e0ad63bf1c4a73ed22fc08
|
|
| BLAKE2b-256 |
3ef845a1e9df10382077705cac6a734d6b2bab9169c632189f9212459a4d9b51
|
Provenance
The following attestation bundles were made for caliper_eval-0.2.2.tar.gz:
Publisher:
publish.yml on edonadei/caliper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
caliper_eval-0.2.2.tar.gz -
Subject digest:
9c74563f8d7daf27ee5d5271f166d852c3a1e19d6aa1c9e16be235d0ff74d191 - Sigstore transparency entry: 1875871742
- Sigstore integration time:
-
Permalink:
edonadei/caliper@c81db8bcec1702a448a911cc8e23642ffd6cbdcd -
Branch / Tag:
refs/tags/v0.2.2 - Owner: https://github.com/edonadei
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@c81db8bcec1702a448a911cc8e23642ffd6cbdcd -
Trigger Event:
release
-
Statement type:
File details
Details for the file caliper_eval-0.2.2-py3-none-any.whl.
File metadata
- Download URL: caliper_eval-0.2.2-py3-none-any.whl
- Upload date:
- Size: 50.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4852daaf423c67f6d4b3eac5a7f74c0d13423f0764ff04ae4872753d2b0ac6c8
|
|
| MD5 |
8d15fc5d59d8e8b385d513b37ab931a1
|
|
| BLAKE2b-256 |
ca4bde40de922a2ef8370c92108631734d8239a415bab70300bae6b2a2a55c0b
|
Provenance
The following attestation bundles were made for caliper_eval-0.2.2-py3-none-any.whl:
Publisher:
publish.yml on edonadei/caliper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
caliper_eval-0.2.2-py3-none-any.whl -
Subject digest:
4852daaf423c67f6d4b3eac5a7f74c0d13423f0764ff04ae4872753d2b0ac6c8 - Sigstore transparency entry: 1875871893
- Sigstore integration time:
-
Permalink:
edonadei/caliper@c81db8bcec1702a448a911cc8e23642ffd6cbdcd -
Branch / Tag:
refs/tags/v0.2.2 - Owner: https://github.com/edonadei
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@c81db8bcec1702a448a911cc8e23642ffd6cbdcd -
Trigger Event:
release
-
Statement type: