CLI for evaluating Claude Code skills and AI agents
Project description
Caliper — Reliability testing for agent skills
Know whether your skill actually works. Write a short spec of what "good" looks like, run it k times, and get a success rate you can track. Caliper also runs the tasks without the skill, so you can see whether it's the skill or the base agent doing the work. Works with the agent you already use: Claude Code, Codex, Pi, or Hermes.
Teach your agent to evaluate:
npx skills@latest add edonadei/caliper
Or run it yourself:
caliper run commit-commands.eval.yaml --k 3 --baseline
You write a spec — a few lines of YAML describing what "working" means, which you hand-write or have /grill-skill generate for you. With --baseline, Caliper runs each task with and without the skill and diffs the two runs task by task:
─────────────────── CALIPER — compare — commit-commands ────────────────────
no skill → with skill · k=3
╭───────────────────────┬─────────────────┬────────┬───────────╮
│ Task │ success │ Δ │ attempts │
├───────────────────────┼─────────────────┼────────┼───────────┤
│ Commits a new feature │ 33.3% → 100.0% │ +66.7% │ ✓✗✗ → ✓✓✓ │
│ Commits a bug fix │ 33.3% → 100.0% │ +66.7% │ ✗✓✗ → ✓✓✓ │
╰───────────────────────┴─────────────────┴────────┴───────────╯
Overall 33.3% → 100.0% Δ (matched) +66.7% ↑
Tokens 290K → 180K Δ -38% (-110K)
Wall 1m 1s → 42s Δ -31% (-19s)
Every task got more reliable, and commit-commands got there on 38% fewer
tokens and 31% less wall-clock time. The attempts column shows every
attempt (✓/✗) before → after, so the bare agent's failures are right there —
not hidden in an average.
Agent skills are hard to test. A skill that works on your machine, on this prompt, today, might fail tomorrow after a model update or a one-line prompt edit. Caliper makes reliability measurable: define what success looks like, run the skill repeatedly, and get a success rate you can track over time.
Use Caliper to answer questions like:
- Did my prompt edit actually improve the skill?
- Is the skill doing the work, or would the base agent pass without it?
- Does it still pass the workflows it passed last week?
- Which agent — Claude Code, Codex, Pi, or Hermes — runs this skill more reliably?
Quick start
Path A — Agentic (let your agent drive)
1. Install the skills
npx skills@latest add edonadei/caliper
2. Generate a spec interactively
In your agent (Claude Code or Codex):
/grill-skill ./my-skill/SKILL.md
grill-skill reads your SKILL.md, interviews you, and writes a 3-task .eval.yaml (happy path, edge case, adversarial).
3. Run and measure
/evaluate-skill run my-skill.eval.yaml --k 3 --baseline
Browse past runs:
/evaluate-skill list
/evaluate-skill report my-skill
Path B — CLI (run it yourself)
1. Install the CLI
pipx install caliper-eval # requires Python 3.10+
2. Write a spec
# my-skill.eval.yaml
skill:
path: ./SKILL.md
tasks:
# Autorater — the LLM judge reads the transcript and decides
- name: Writes a conventional commit message
prompt: "Summarize the staged git diff as a commit message."
expect: >
The response is a conventional-commit message: a concise subject
line under 72 characters, followed by a body explaining why the
change was made, not just what changed.
# Script execution — a deterministic Python assertion
- name: Generates a valid config file
cleanup: rm -f /tmp/app.config.json
prompt: "Generate a config at /tmp/app.config.json with a 'port' of 8080."
assert: |
import json
from pathlib import Path
data = json.loads(Path("/tmp/app.config.json").read_text())
assert data["port"] == 8080
expect: is graded by the judge LLM; assert: runs locally as Python. Use either or both.
The spec never names an engine — the skill and judge default to claude-code, and you pick a different agent/model at run time with --model / --judge-model (see Choosing an engine).
3. Run it
caliper run my-skill.eval.yaml --k 3 # add --baseline to diff vs the bare agent
4. Read the output
─────────── CALIPER — my-skill (claude-code) — 2026-06-19 14:23 ───────────
╭──────────────────────────────────┬───────┬─────────┬────────┬──────┬───────────╮
│ Task │ k (3) │ success │ Tokens │ Wall │ │
├──────────────────────────────────┼───────┼─────────┼────────┼──────┼───────────┤
│ Writes a conventional commit msg │ 3/3 │ 100.0% │ 79K │ 27s │ ✓ PASS │
│ Generates a valid config file │ 2/3 │ 66.7% │ 82K │ 33s │ ~ PARTIAL │
╰──────────────────────────────────┴───────┴─────────┴────────┴──────┴───────────╯
Score 83.3% █████████████████░░░
Tokens 159K in / 2K out
Wall 1m 0s 10.0s per attempt
Results saved to .caliper/results/my-skill/2026-06-19T14-23-01Z.json
--verbose adds pass@k and pass^k columns (both derived from the raw rate).
Not sure what to put in a spec?
The Eval Starter Pack has four copy-paste templates, each catching a real agent failure (false success, tool misuse, runaway loops, prompt regressions). Every template runs green as-is against a bundled example, then points at your own skill by editing two or three commented lines.
How it works
.eval.yaml spec
│
▼
Harness ──── runs your skill against the agent (Claude Code / Codex / Pi)
│
▼
Judge ──── LLM autorater and/or deterministic Python assertions
│
▼
success rate + saved transcript
Each attempt runs in an isolated temporary home with no session history. Results are saved as JSON you can inspect and diff later.
Agent skills
The repo ships two agent skills. Install both with:
npx skills@latest add edonadei/caliper
evaluate-skill — run and manage evals
Create, validate, run, and summarize evals from inside your normal workflow — no separate terminal needed. The skill installs Caliper automatically if it's missing.
Then use it in Claude Code:
/evaluate-skill run my-skill.eval.yaml --k 3
/evaluate-skill validate my-skill.eval.yaml
Or in Codex:
Use the evaluate-skill skill to run my-skill.eval.yaml with k=3 and summarize the result.
grill-skill — create evals interactively
Don't have evals yet? grill-skill guides you through creating them. It reads your SKILL.md, interviews you about what good behavior looks like, and generates a 3-task spec (happy path, edge case, adversarial). Then it runs the eval and loops — k=1 to validate, k=3 to measure, baseline before you commit.
/grill-skill ./my-skill/SKILL.md
No path needed if you're already in the skill's directory:
/grill-skill
If an .eval.yaml already exists next to your skill, grill-skill reads the existing tasks and interviews you about gaps instead of starting from scratch.
Core concepts
| Term | What it is |
|---|---|
| Spec | A .eval.yaml file that describes the skill, judge, and tasks to run |
| Backend | The CLI agent that executes the skill (claude-code, codex, pi, hermes) |
| Judge | What decides pass/fail — an LLM reading the transcript (expect:), Python assertions (assert:), or both |
| success rate | The primary score: run k times, measure how often a single run works (pass@k/pass^k are secondary views, under --verbose) |
| Baseline | Re-run the same tasks without the skill to prove the skill is doing the work |
| Attempt | One isolated run of a single task — fresh temporary home, no session history |
Choosing an engine
The engine (backend + model) is a runtime axis, not a spec field — the spec
describes what is tested and how success is judged, and you pick the agent
that runs and grades it at invocation. Both default to claude-code; select a
different one with --model / --judge-model:
caliper run my-skill.eval.yaml # claude-code (default)
caliper run my-skill.eval.yaml --model codex # codex, its default model
caliper run my-skill.eval.yaml --model codex:gpt-5-codex
caliper run my-skill.eval.yaml --model pi --judge-model claude-code
| Backend | Requires | Best for |
|---|---|---|
claude-code |
Claude Code CLI installed and authenticated | Testing Claude Code slash-command skills |
codex |
Codex CLI installed (npm install -g @openai/codex) |
Testing Codex skills |
pi |
pi CLI installed (npm install -g @earendil-works/pi-coding-agent) and authenticated |
Testing pi skills (agentskills.io) |
hermes |
Hermes Agent CLI installed and authenticated (Nous Research) | Testing skills on Hermes; hermes:<provider>/<model> selects the model |
Caliper runs skills only through CLI agents — every backend can actually load and run a skill. There is no direct-API backend: to run against API-priced billing, configure one of these CLIs with an API key (e.g. ANTHROPIC_API_KEY / OPENAI_API_KEY) rather than selecting a separate backend.
The skill engine and judge engine are independent — you can test a Codex skill with a Claude judge, or any other combination, by pairing --model with --judge-model.
Claude Code setup
Install and authenticate the claude CLI. --model claude-code uses your existing Claude Code auth — no extra configuration needed.
Codex setup
npm install -g @openai/codex
codex login
--model codex calls codex exec. If the Codex desktop app is installed, Caliper prefers the app-bundled binary over codex on PATH. Set CODEX_CLI_PATH to force a specific binary.
pi setup
npm install -g @earendil-works/pi-coding-agent
pi # then authenticate (e.g. /login for a subscription provider, or set the provider API key)
--model pi runs pi --print --mode json and loads the skill natively via pi's --skill flag (the agentskills.io standard). It reuses your ~/.pi/agent auth and settings — the :model half of --model pi:<model> overrides pi's configured default when set. Set PI_CLI_PATH to force a specific binary. Note: pi's built-in default provider is google, so running --model pi with no model relies on your pi config to resolve a provider you are authenticated for.
Hermes setup
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes login # authenticate; pick a default model/provider you have credits for
Hermes is a stateful, always-on agent (persistent memory, a persona, auto-generated skills), so Caliper normalizes it to a neutral agent to keep its score apples-to-apples with the other backends: every attempt runs in an isolated HERMES_HOME seeded with your ~/.hermes auth/config only (never SOUL.md/MEMORY.md) and with --ignore-rules, and the skill-under-test is staged as the sole local skill. --model hermes runs hermes -z (oneshot) then hermes sessions export to recover the full tool-call trajectory; --model hermes:<provider>/<model> (e.g. hermes:anthropic/claude-sonnet-4.6) selects the model, otherwise your ~/.hermes/config.yaml default is used — point it at a provider you have credits for. Set HERMES_CLI_PATH to force a specific binary. Hermes updates itself (hermes update), so it is not part of caliper update-cli.
Check installed CLI versions:
caliper update-cli --check
Recommended workflow
- Create a spec for one behavior you care about.
- Run with
--k 1while iterating on the spec. - Add
assert:for facts an LLM judge might guess wrong (files, JSON, command output). - Move to
--k 3or higher once the task is stable. - Add
--baselineto prove the skill is making a difference. - Commit the spec alongside the skill so contributors can run the same eval.
/evaluate-skill run my-skill.eval.yaml --k 3 --baseline --verbose
Spec format
To scaffold a spec, use the evaluate-skill
or grill-skill skill, or hand-write
the YAML below.
skill:
path: ./SKILL.md # path to the skill file (optional for baseline-only runs)
# Note: there is no `backend`/`model` or `judge:` block. The engine is a runtime
# axis — pass `--model` / `--judge-model` at run time (default: claude-code).
sandbox:
extra_path:
- ./bin # prepended to PATH inside each attempt
forbidden_files:
- ".*\\.eval\\.yaml$" # prevents agent from reading the spec
- "./.caliper/.*" # prevents agent from reading saved results
tasks:
- name: Short task name
setup: <shell command> # optional, runs before each attempt
cleanup: <shell command> # optional, always runs after each attempt
prompt: <prompt sent to the agent>
expect: <natural-language success condition>
assert: |
# optional inline Python assertion
assert True
- name: Task with external assertion script
prompt: "Generate a report"
assert: ./assertions/check_report.py
Each task needs at least one of expect or assert. Task IDs are assigned automatically as task-001, task-002, and so on.
Judging
LLM autorater (expect:)
The judge engine reads the full attempt transcript and decides whether the expect condition was met. When the backend captures tool-call traces (Claude Code, Codex, pi, Hermes), those traces are included — the judge can verify things like "the agent used tool X" without relying on the final text alone.
The judge engine is chosen at run time and defaults to claude-code; point it at a different agent with --judge-model (e.g. --judge-model codex), independently of the skill's --model.
Deterministic assertions (assert:)
Python assertions run locally. Use these for facts the LLM judge might guess:
- file exists / exact file contents
- JSON / schema validity
- command output
- images or screenshots
- repository state
tasks:
- name: Writes an output file
cleanup: rm -f /tmp/out.txt
prompt: "Write hello world to /tmp/out.txt"
assert: |
from pathlib import Path
path = Path("/tmp/out.txt")
assert path.exists(), "Output file was not created"
assert path.read_text().strip() == "hello world"
When both expect and assert are present, both must pass.
CLI reference
| Command | Description |
|---|---|
caliper run <spec> |
Run an evaluation spec |
caliper validate <spec> |
Validate a spec file |
caliper list [spec] |
List specs and saved runs |
caliper report <spec-or-result> |
Re-render saved results |
caliper compare <A> <B> |
Diff two saved runs of the same eval, task by task |
caliper update-cli [backend] |
Check or update installed agent CLI versions |
caliper run flags
| Flag | Default | Description |
|---|---|---|
--k INT |
3 |
Attempts per task |
--baseline |
off | Also run each task without the skill |
--workers INT |
4 |
Parallel task workers |
--timeout INT |
120 |
Seconds per attempt |
--fail-fast INT |
0 |
Stop a task after N consecutive infra_error/timeout attempts (0 disables) |
--model TARGET |
claude-code |
Skill engine — backend and/or model (see below) |
--judge-model TARGET |
claude-code |
Judge engine — backend and/or model (see below) |
--verbose |
off | Show per-attempt judge reasoning |
--output PATH |
— | Also save results JSON to a specific path |
--model and --judge-model syntax
The engine is not stored in the spec — these flags select it, defaulting to claude-code when omitted. Both accept a backend:model compound value, a bare backend name, or a bare model name:
# Backend and model together
caliper run my-skill.eval.yaml --model codex:gpt-5-codex
# Backend only (that backend's default model)
caliper run my-skill.eval.yaml --model codex
# Model only (backend stays claude-code)
caliper run my-skill.eval.yaml --model claude-sonnet-4-6
# Select the judge engine independently
caliper run my-skill.eval.yaml --model codex --judge-model claude-code:claude-haiku-4-5-20251001
Accepted backends: claude-code, codex, pi, hermes (alias: claude → claude-code). The actual engine used is recorded in each saved run's RunMeta — the skill backend/model, and the judge_backend/judge_model that graded it — so results stay traceable even though the spec doesn't pin it. When you don't name a model and the CLI uses its own default, RunMeta records the concrete model the agent resolved rather than a bare "default", wherever the backend reports it — the skill model from hermes' session export, and the judge_model from the claude-code judge's JSON output. judge_model stays empty for an assert:-only run, where no LLM judge fired.
Comparing two runs (caliper compare)
An ablation compares two runs of the same eval — a full skill vs. a
shortened variant, or the same skill over time. caliper compare <A> <B> diffs
two already-saved runs task by task, so you don't hand-write a JSON script to
answer "did this change regress?".
# Latest run of each spec (a bare spec name resolves to its latest run)
caliper compare commit-simple-full commit-simple-short
# Pin specific runs by pointing at their results JSON
caliper compare .caliper/results/demo/2026-07-01T10-00-00Z.json \
.caliper/results/demo/2026-07-02T09-00-00Z.json
# Machine-readable diff for a ship / no-ship decision
caliper compare A B --format json
Each positional (A, B) is addressed exactly like report's argument: a spec
name (→ its latest run) or a path to a results JSON. There are no --run-a/-b
flags — pin a historical run by naming its JSON path.
──────────────────── CALIPER — compare — commit-simple ─────────────────────
2026-07-01T10-00-00Z (claude-code) → 2026-07-02T09-00-00Z (claude-code) · k=5
╭──────────────────┬─────────────────┬────────┬───────────────╮
│ Task │ success │ Δ │ attempts │
├──────────────────┼─────────────────┼────────┼───────────────┤
│ commits cleanly │ 100.0% → 100.0% │ — │ ✓✓✓✓✓ → ✓✓✓✓✓ │
│ handles conflict │ 100.0% → 20.0% │ -80.0% │ ✓✓✓✓✓ → ✓✗✗✗✗ │
│ pushes upstream │ 80.0% → — │ — │ ✓✓✓✓✗ → ⊘⊘⊘⊘⊘ │
╰──────────────────┴─────────────────┴────────┴───────────────╯
Overall 100.0% → 60.0% Δ (matched) -40.0% ↓
Tokens 1.2M → 700K Δ -42% (-500K)
Wall 6m 18s → 3m 40s Δ -42% (-2m 38s)
⚠ 1 regression: handles conflict
⊘ 1 unmeasured (excluded from Δ): pushes upstream
unmatched — only in A: flaky task only in B: new task
How the diff reads:
- Each row reads
before → after. The two runs are named once in the header (here, the two timestamps; a--baselinerun saysno skill → with skill), so there's no A/B legend to decode. - Tasks are matched by name.
task_idis only positional, so name is the stable identity — reordered tasks still line up. A task present in only one run is listed as unmatched and left out of the delta. Δisafter − before. A negative Δ renders red and flags the task as a regression (any-below rule: the after side below the before side).- The score excludes unusable attempts. The
attemptscolumn reuses the run report's glyphs;⊘marks an unusable attempt (rate-limit / timeout / judge error). A side with no usable attempts shows—(unmeasured) and is never counted as a regression — infra noise can't fake a loss. - The headline
Δ (matched)averages each side over only the tasks measured on both sides, so it is strictly like-for-like. - Token and wall-clock deltas sit under the headline as secondary signals:
a drop renders green (cheaper), a rise red (costlier). They are never a
regression — a token drop at a flat score is the win an ablation is looking for,
and a rise is a trade-off to weigh, not a failure. Only the score feeds
has_regression. The token row is shown only when both runs reported tokens; wall time is always shown. Dollar cost is deliberately not tracked (tokens are the volume signal). - Guards for a
kmismatch or different spec names print as warnings in the header and in--format json(k_mismatch,spec_mismatch,warnings), so an agent drivingcomparesees them too.
The --format json output serializes the full comparison (per-task scores,
deltas, regression/has_regression flags, unmatched task lists, the warnings,
and per-side usage totals a_usage/b_usage with token_delta/wall_delta)
for scripting.
Scoring
Every attempt carries a typed outcome, so infrastructure and judge noise are not scored as task failure:
| Outcome | Meaning | Counts toward the score? |
|---|---|---|
pass |
satisfied the task's judge(s) | ✅ success |
task_fail |
the skill genuinely failed the task | ✅ attempt |
cheat |
a forbidden-file read was detected | ✅ attempt |
infra_error |
harness failure — nonzero exit, or a detected rate-limit / spending-cap | ❌ unusable |
timeout |
exceeded the time budget with no result | ❌ unusable |
judge_error |
the judge produced no verdict (unparseable / errored autorater) | ❌ unusable |
passed is retained in the JSON as a convenience, equal to outcome == pass.
The primary metric is the raw success rate — how often a single run works, computed over the usable attempts (those that got a fair shot); unusable attempts leave the denominator and are reported as a separate "N unusable" count:
usable = pass + task_fail + cheat
score = successes / usable # raw rate; None if usable == 0
Two secondary views are kept for anyone who wants them (shown under --verbose,
and on every task in the JSON as pass_at_k / pass_hat_k):
pass@k = 1 - (1 - score) ^ usable # P(≥1 of k passes)
pass^k = score ^ usable # P(all k pass)
Which one to look at depends on how the skill is actually used:
| The question you're asking | Metric | For a 1/3 skill (k=3) |
|---|---|---|
| How reliable is a single run? (default) | success rate | 33% |
| If I retry up to k times and keep any win, do I get one? | pass@k |
70% |
| Will it work on every run, no exceptions? | pass^k |
4% |
Reach for pass@k when a failure is cheap to retry and you can keep the good
run (a human re-runs it, or you regenerate until it works) — it's the
retry-optimistic, "eventual success" view, always ≥ the raw rate. Reach for
pass^k when the skill runs unattended or as one link in a chain, where a
single failure breaks the whole thing — it's the strict "must never fail" view,
always ≤ the raw rate. When in doubt, the success rate is the honest
middle: pass@k is the code-generation metric (generate k, keep the best) and
flatters flaky skills (1/3 → 70.4%), which is why Caliper leads with the rate.
The aggregate is the average task success rate, skipping tasks with no usable
attempts. With --baseline, Caliper runs the same tasks without the skill and
reports the delta. A throttled or judge-flaked run no longer masquerades as a
regression.
For persistent infrastructure failures, --fail-fast N can stop scheduling new
attempts for a task after N consecutive infra_error or timeout outcomes.
The default 0 keeps the historical behavior and runs all k attempts. An
early-stopped task is shown as ABORTED in the report; if every completed
attempt was unusable, its score remains null and it is skipped in the
aggregate score.
Token & wall-clock usage
Pass@k tells you whether a skill works; usage tells you what it costs to get there. Two runs can have identical scores while one burns twice the tokens. Caliper records token volume and wall-clock time per attempt and rolls them up per run:
With skill 100.0% ████████████████████
Tokens 1.2M in / 340K out
Wall 6m 18s 12.6s per attempt
⊘ unusable spend: 180K tokens, 42s (2 attempts, not counted in the average)
- The results table carries per-task
Tokens(total) andWallcolumns, so you can see which task is the expensive one at a glance; the run summary line below it aggregates across the whole run. - Each
AttemptRecordcarries an optionalusageobject withinput_tokens(non-cached prompt),output_tokens,cache_read_tokens,cache_creation_tokens, and a computedtotal_tokens. The four token fields are disjoint —input_tokensexcludes cache — sototal_tokensnever double-counts.duration_seconds(already recorded) is the wall-clock half. in= input + cache_read + cache_creation;out= output. Every attempt counts toward the run total; the unusable slice (timeout / infra / judge error) is broken out separately so wasted spend is visible without distorting the per-usable-attempt average.- All usage fields are optional — a backend that can't report them leaves
them
nulland the line renders—. Support:claude-code,codex,pi, andhermesall report token usage;codexuses OpenAI semantics (itsinput_tokensincludes cache) so it is normalized to the non-cached contract. - Dollar cost is deliberately not tracked — it is inconsistent across backends and would need a maintained price table. Tokens are the volume signal; a dollar figure can be derived downstream if needed.
- With
--baseline, the whole no-skill run is retained (RunResults.baseline_task_results) and the report renders as acompareview — the same table, attempt strips, and token/wall deltas — so the skill-vs-bare- agent difference is shown side by side (see the quickstart at the top). report --format jsonincludes a derivedusage_totalsblock; the saved results JSON keeps the raw per-attemptusage(totals are derived, never persisted), plus the fullbaseline_task_resultswhen--baselineran.comparesurfaces token/wall deltas across two runs — see above.
Project layout
caliper/
commands/ CLI command implementations
harness/ Agent execution backends (Claude Code, Codex, pi, Hermes)
judge/ LLM and script judging implementations
schema/ Eval spec and result models
runner.py Evaluation orchestration
skills/
evaluate-skill/ Agent skill for running Caliper from Claude Code or Codex
grill-skill/ Agent skill for creating and iterating on evals interactively
tests/ Pytest coverage for harnesses, judges, and runner behavior
Contributing
Contributions are welcome. See CONTRIBUTING.md for good first areas, the pre-PR checklist, the ruff formatting convention and pinned version, and the one-time pre-commit install step.
Troubleshooting
codex judge failed: model ... is not supported
The model name is not available to your Codex account. Use a model that codex exec --model <name> accepts.
codex CLI not found
npm install -g @openai/codex
claude command not found
Install and authenticate Claude Code, or switch the backend to codex or pi.
A task passes only because of assert:
When a task has only assert:, no LLM judge runs. Add expect: if you also want an LLM to evaluate the transcript.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file caliper_eval-0.8.0.tar.gz.
File metadata
- Download URL: caliper_eval-0.8.0.tar.gz
- Upload date:
- Size: 164.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
27d4cb9f25d47f65279a17a48d8a28aed3abca207c46e38ebd4ba6913dad9080
|
|
| MD5 |
dad7a819e80cec723e185a45c3f88cce
|
|
| BLAKE2b-256 |
1fae66712497c7049ef53adb6f0471c4ed87b51df9a8f70fd6998b750b161883
|
Provenance
The following attestation bundles were made for caliper_eval-0.8.0.tar.gz:
Publisher:
publish.yml on edonadei/caliper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
caliper_eval-0.8.0.tar.gz -
Subject digest:
27d4cb9f25d47f65279a17a48d8a28aed3abca207c46e38ebd4ba6913dad9080 - Sigstore transparency entry: 2082759564
- Sigstore integration time:
-
Permalink:
edonadei/caliper@58d1af5ecf0386a9b22c66f8c34a52cf12e188c4 -
Branch / Tag:
refs/tags/v0.8.0 - Owner: https://github.com/edonadei
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@58d1af5ecf0386a9b22c66f8c34a52cf12e188c4 -
Trigger Event:
release
-
Statement type:
File details
Details for the file caliper_eval-0.8.0-py3-none-any.whl.
File metadata
- Download URL: caliper_eval-0.8.0-py3-none-any.whl
- Upload date:
- Size: 74.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7c8cc783be0738ae6ba68cc0abf3dad5cb90d8070499bc725a51356e9d192ebc
|
|
| MD5 |
c22c53e25bc414ac338ee71d9116b7ac
|
|
| BLAKE2b-256 |
4915a2c89331935bdbdb3147e2ce7e4cef02fc05689a8e76705a344c9d00172d
|
Provenance
The following attestation bundles were made for caliper_eval-0.8.0-py3-none-any.whl:
Publisher:
publish.yml on edonadei/caliper
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
caliper_eval-0.8.0-py3-none-any.whl -
Subject digest:
7c8cc783be0738ae6ba68cc0abf3dad5cb90d8070499bc725a51356e9d192ebc - Sigstore transparency entry: 2082759573
- Sigstore integration time:
-
Permalink:
edonadei/caliper@58d1af5ecf0386a9b22c66f8c34a52cf12e188c4 -
Branch / Tag:
refs/tags/v0.8.0 - Owner: https://github.com/edonadei
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@58d1af5ecf0386a9b22c66f8c34a52cf12e188c4 -
Trigger Event:
release
-
Statement type: