evalkit
Evals as code for LLM apps and agents. Write tasks and graders in YAML, run them in CI, and fail the build when quality regresses.
You describe what good output looks like in a YAML file. evalkit sends each task to a model or to your own app, grades the answers, and gives you a results table, JSON and JUnit output, a regression diff against a saved baseline, and a badge. It runs the same on your laptop and in GitHub Actions.
$ evalkit run examples/support-bot/evals.yaml
evalkit 0.2.0 suite support-bot target command:python3 8 tasks x 1 run
STATUS TASK SCORE RUNS LATENCY COST DETAIL
pass password-reset 1.00 1/1 40ms -
pass order-status 1.00 1/1 30ms -
pass unknown-order 1.00 1/1 20ms -
pass refund 1.00 1/1 41ms -
pass outage 1.00 1/1 23ms -
pass handoff 1.00 1/1 23ms -
pass pii-echo 1.00 1/1 20ms -
FAIL cancel-subscription 0.67 0/1 19ms - contains: missing 'Billing > Subscription'
7/8 passed (87.5%), score 0.96, 0 flaky, 0 errored, latency p50 23ms p95 41ms, 0/8 runs cached, 0.11s
PASS: pass rate 87.5% meets fail_under 85.0%
That output is real. The example's target is a small rules-based bot that runs as a command, with one known gap, so the suite passes at its 85% bar with one failing task.
Install
Install from PyPI into a Python 3.11+ environment. The
distribution is named sic-evalkit. The command and the import package are both evalkit:
pip install sic-evalkit
evalkit --version
To install a standalone executable instead, run the installer. It picks the right release
asset for your OS and CPU, checks it against SHA256SUMS, and puts evalkit in
~/.local/bin (set EVALKIT_INSTALL_DIR to change that, or EVALKIT_VERSION=v0.2.0 to pin a
version):
curl -fsSL https://raw.githubusercontent.com/superintelligenceco/evalkit/main/install.sh | sh
Every GitHub Release carries a
standalone evalkit executable for each platform, the Python wheel and sdist, and a
SHA256SUMS file. The executables bundle Python, so you don't need Python installed.
| Platform | Asset |
|---|---|
| Linux x64 (glibc 2.31 or later) | evalkit-linux-x64 |
| Linux arm64 (glibc 2.31 or later) | evalkit-linux-arm64 |
| macOS on Apple silicon | evalkit-macos-arm64 |
| macOS on Intel | evalkit-macos-x64 |
| Windows x64 | evalkit-windows-x64.exe |
To download an executable by hand, check it, and put it on your PATH:
curl -fLO https://github.com/superintelligenceco/evalkit/releases/latest/download/evalkit-linux-x64
curl -fLO https://github.com/superintelligenceco/evalkit/releases/latest/download/SHA256SUMS
sha256sum --check --ignore-missing SHA256SUMS
chmod +x evalkit-linux-x64
sudo mv evalkit-linux-x64 /usr/local/bin/evalkit
evalkit --version
On macOS, use shasum -a 256 --check --ignore-missing SHA256SUMS. The macOS and Windows
executables aren't signed, so a browser download can trigger Gatekeeper or SmartScreen. Download
with curl instead, or on macOS run xattr -d com.apple.quarantine evalkit-macos-arm64.
Container image
Each release also publishes a multi-arch (amd64 and arm64) image to
ghcr.io/superintelligenceco/evalkit. Its working directory holds the starter suite, so this runs
offline:
docker run --rm ghcr.io/superintelligenceco/evalkit run evals.yaml
Mount your project on /work to run your own suites:
docker run --rm -u "$(id -u):$(id -g)" -v "$PWD:/work" ghcr.io/superintelligenceco/evalkit run evals.yaml
Tags are latest, the version (such as 0.2.0), and edge for the tip of main. The images are
signed with cosign and carry build provenance attestations.
What the executables include
The executables run everything evalkit does, including command and http targets and python
graders. A python grader file runs on the bundled Python 3.12. It can import evalkit's own
dependencies (httpx, jsonschema, pydantic, yaml) and the standard library modules that
the executable bundles, which cover common ones such as re, json, csv, and statistics but
not all of them. If your graders import other packages, install evalkit with pip into an
environment that has them.
Quickstart
These commands run offline. The starter suite uses the built-in mock provider, so you
don't need an API key.
evalkit init
evalkit run evals.yaml
To install from source instead, run pip install "git+https://github.com/superintelligenceco/evalkit".
To test a real model, replace the target in evals.yaml:
target:
provider: openai # any OpenAI-compatible API
model: gpt-4o-mini # reads OPENAI_API_KEY
Why evalkit
Prompt and model changes break LLM apps quietly. A reworded system prompt, a new model version, or a retrieval tweak can drop quality on cases that worked last week, and nothing fails until a user notices. Unit tests don't catch it, because the output isn't deterministic and "correct" is often a judgment call.
evalkit treats evals the way you already treat tests:
- They live in the repo. A suite is a reviewable YAML file next to your code.
- They run in CI. The exit code, JUnit XML, and job summary plug into any pipeline.
- They gate merges. Compare a run against a saved baseline and fail on a real drop, not noise.
- They're cheap to rerun. Responses are cached in SQLite, so a rerun with unchanged prompts makes no API calls.
- They test your app, not only a model. Point the target at a command or an HTTP endpoint and grade what your users actually see.
Features
- Graders:
exact,contains,regex,json_schema,numericwith tolerances, custompythonfunctions, andllm_judgewith a rubric and any provider as the judge. Any grader can be negated and weighted. - Targets and providers:
openai(OpenAI and any compatible gateway or local server),anthropic,command(your program, prompt on stdin),http(your endpoint, templated JSON body), and a deterministicmockprovider for offline tests and demos. - Runner: bounded concurrency, retries with exponential backoff on timeouts, 429, and 5xx, and a SQLite response cache.
- Flakiness stats: run each task N times with
repeat. evalkit reports per-task pass rates and marks tasks that pass on some runs and fail on others as flaky. - Cost and latency: per-task token counts, cost from prices you supply, and p50/p95 latency.
- Outputs: terminal table, full results JSON, JUnit XML, GitHub-flavored Markdown, and a shields.io endpoint badge.
- Regression gate:
evalkit compare base.json head.jsonfails when the pass rate or mean score drops beyond a threshold, or, with--strict, when any single task regresses. - GitHub Action: runs a suite, writes the Markdown summary to the job summary, and exposes the pass rate as step outputs.
- Editor support:
evalkit schemaprints a JSON Schema for suite files.
How it works
flowchart LR
Y[evals.yaml] --> R[runner]
R -->|prompt| T[target: model, command, or HTTP app]
T -->|output| G[graders]
G -->|llm_judge| J[judge model]
R <--> C[(SQLite cache)]
G --> O[table, JSON, JUnit, Markdown, badge]
O --> D[compare with baseline]
Examples
| Suite | What it shows | Needs a key |
|---|---|---|
examples/quickstart |
One task per deterministic grader, on the mock provider. |
No |
examples/support-bot |
A command target, tasks from JSONL, python graders, an llm_judge. |
No |
examples/flaky |
repeat: 10 and per-task pass rates to separate flaky tasks from broken ones. |
No |
examples/openai-compatible |
A real model plus a judge over any OpenAI-compatible API. | Yes, or a local server |
Flakiness stats come from repeating each task:
$ evalkit run examples/flaky/evals.yaml
evalkit 0.2.0 suite flakiness target mock:sometimes-wrong 3 tasks x 10 runs
STATUS TASK SCORE RUNS LATENCY COST DETAIL
pass planet 1.00 10/10 0ms -
pass gold 0.70 7/10 0ms -
FLAKY hamlet 0.40 4/10 0ms - contains: missing 'Shakespeare'
2/3 passed (66.7%), score 0.70, 2 flaky, 0 errored, latency p50 0ms p95 0ms, 0/30 runs cached, 0.05s
PASS: pass rate 66.7% meets fail_under 66.0%
gold passes because 7 of 10 runs meet the suite's task_pass_rate: 0.7. It still counts toward
the flaky total, since its runs disagree.
Catch regressions
Save a baseline on your main branch, then compare every change against it:
evalkit run evals.yaml -o base.json # on main
evalkit run evals.yaml -o head.json # on your branch
evalkit compare base.json head.json --threshold 0.05
Here is real output after breaking two replies in the support-bot example:
$ evalkit compare base.json head.json
pass rate 87.5% -> 62.5% (-25.0 pts)
score 0.958 -> 0.854 (-0.104)
regressed (2):
order-status: pass -> fail (score 1.00 -> 0.50)
refund: pass -> fail (score 1.00 -> 0.67)
REGRESSION: pass rate dropped 25.0 points (threshold 0.0)
REGRESSION: score dropped 0.104 (threshold 0.000)
compare exits with status 1 on a regression. You can also gate in one step with
evalkit run evals.yaml --baseline base.json --threshold 0.05.
Suite file reference
A suite is a YAML (or JSON) file. Unknown keys are errors, so typos fail loudly. Run
evalkit validate evals.yaml to check a file without running it, and evalkit schema for a JSON
Schema your editor can use.
version: 1 # optional, the only version is 1
name: support-bot # required
description: Optional text shown in reports.
target: {provider: openai, model: gpt-4o-mini} # required, see "Providers"
judge: {provider: openai, model: gpt-4o-mini} # default judge for llm_judge graders
system: You are a helpful support agent. # system prompt for the target
prompt: "Customer says: {{message}}" # template; omit it to send each task's input as-is
graders: # applied to every task, before each task's own graders
- type: contains
value: ["Thanks"]
tasks: # a list, or a path to a .yaml, .json, or .jsonl file
- id: refund # defaults to task-1, task-2, ...
vars: {message: "I want a refund"}
input: null # string prompt, or any value when you use `prompt`
expected: Billing > Refunds # default value for exact, contains, and numeric
tags: [billing] # select with `evalkit run -t billing`
metadata: {owner: payments} # free-form, passed to python graders
graders:
- type: regex
pattern: 'Billing\s*>\s*Refunds'
settings:
concurrency: 4 # parallel requests
repeat: 1 # runs per task, for flakiness stats
retries: 2 # retries on timeouts, 429, and 5xx
retry_backoff: 0.5 # seconds, doubles per attempt
seed: 0 # passed to providers that accept a seed
cache: true # cache model responses in SQLite
cache_path: .evalkit/cache.sqlite # relative to the directory you run evalkit from
fail_under: 1.0 # minimum share of passing tasks for exit code 0
task_pass_rate: 1.0 # share of a task's repeats that must pass
Templates
Prompts, grader values, regex patterns, rubrics, and http bodies accept {{ name }}
placeholders. The variables are input, expected, id, metadata, vars, and each key of
vars (and of input, when it's a mapping) at the top level. Dotted paths such as
{{ vars.city }} and {{ items.0 }} walk into nested values. A value that is exactly one
placeholder keeps its type, so value: "{{ expected }}" can resolve to a number.
Provider settings in target and judge also expand environment variables: ${VAR} or
${VAR:-default}. Task data is never expanded.
Providers
Every provider accepts timeout (seconds, default 60), cache (override the default), and
pricing: {input_per_mtok, output_per_mtok} in USD per million tokens. evalkit ships no price
table, so cost is reported only when you set prices.
| Provider | Key fields | Cached by default |
|---|---|---|
openai |
model, base_url (default https://api.openai.com/v1), api_key_env (default OPENAI_API_KEY, null for none), temperature, max_tokens, headers, params |
Yes |
anthropic |
model, base_url, api_key_env (default ANTHROPIC_API_KEY), anthropic_version, temperature, max_tokens (default 1024), headers, params |
Yes |
command |
command (string or list), cwd, env. The prompt goes to stdin unless an argument contains {{prompt}}. stdout is the output. EVALKIT_SEED, EVALKIT_REPEAT, and EVALKIT_SYSTEM are set. |
No |
http |
url, method (POST, PUT, or GET), headers, body (JSON template, default {"input": "{{prompt}}"}), output_path (dotted path into the response, such as data.answer) |
No |
mock |
responses (list of {match: regex, output, flaky}), default (echoes the prompt if unset), latency_ms, flaky, flaky_output, fail_times |
Yes |
The openai provider works with any server that speaks the chat completions API, including
gateways such as LiteLLM and OpenRouter and local servers such as vLLM, llama.cpp, and Ollama.
Set base_url and, for servers without auth, api_key_env: null.
params merges extra fields into the request body, for example params: {top_p: 0.9}.
Scoring
Each grader returns pass or fail, a score from 0 to 1, and a reason. A run passes when every
grader passes. Its score is the weighted mean of the grader scores. A task passes when at least
task_pass_rate of its runs pass. The suite passes when at least fail_under of its tasks pass.
Task statuses in reports are pass, FAIL, FLAKY (some runs passed, too few to pass the task),
and ERROR (the target or a grader raised an error).
Graders reference
Every grader accepts name (label in reports), negate: true (invert the verdict, for "must not"
checks), and weight (default 1, used in the task score).
| Type | Passes when | Options |
|---|---|---|
exact |
The output equals value. |
value (defaults to expected), ignore_case, strip (default true) |
contains |
The output contains value, or each item of a list. |
value, mode: all or any, ignore_case. Score is the share of items found. |
regex |
pattern matches anywhere in the output. |
pattern, ignore_case, fullmatch (match the whole stripped output) |
json_schema |
The output is JSON that validates against the schema. | schema (inline) or schema_file (JSON or YAML, relative to the suite), extract (default true: find JSON in code fences or prose) |
numeric |
A number in the output is within tolerance of value. |
value (defaults to expected), abs_tol, rel_tol, pick: last or first. Commas such as 1,250.50 are handled. |
python |
Your function says so. | function (checks.py:name relative to the suite, or package.module:name), args (keyword arguments) |
llm_judge |
A judge model scores the output at or above min_score. |
rubric (templated), provider (defaults to the suite judge), scale (default 5), min_score (default scale - 1) |
Python graders
A python grader receives the output and a context dict with input, expected, vars, id,
metadata, prompt, repeat, and seed. It can be sync or async and returns one of:
def within_length(output, context, max_words=40):
words = len(output.split())
return {
"pass": words <= max_words,
"score": min(1.0, max_words / max(words, 1)),
"reason": f"{words} words (max {max_words})",
}
# Also valid: True / False, a score in [0, 1] (passes at 0.5), or (passed, "reason").
LLM-as-judge
The judge sees the rubric, the task input, the task's expected value as a reference when there
is one, and the output. It replies with {"score": <1..scale>, "reason": "..."}. evalkit
normalizes the score to 0 to 1 and passes the grade at min_score. Judge calls go through the same
retries and cache as the target. A reply without a readable score fails the grade and shows the raw
reply in the reason.
judge:
provider: anthropic # reads ANTHROPIC_API_KEY
model: ${JUDGE_MODEL}
graders:
- type: llm_judge
name: tone
rubric: The reply is polite, acknowledges the customer, and gives a concrete next step.
min_score: 4
Command-line reference
| Command | What it does |
|---|---|
evalkit run SUITE |
Run a suite. Writes outputs with -o results.json, --junit, --markdown, --badge. Filter with -t TAG, -k TEXT, --limit N. Override -j, -n (repeat), --seed, --no-cache, --fail-under. Gate with --baseline, --threshold, --strict. |
evalkit compare BASE HEAD |
Diff two results files. --threshold 0.05 allows a 5-point drop. --strict fails on any regressed task. --format text, markdown, or json. |
evalkit validate SUITE... |
Check suite files without running them. |
evalkit init [DIR] |
Write a starter evals.yaml that passes offline. |
evalkit schema |
Print the suite JSON Schema. |
evalkit cache stats or clear |
Inspect or empty the response cache. |
Exit codes: 0 passed, 1 failed the pass-rate bar or regressed, 2 usage or config error.
NO_COLOR and FORCE_COLOR control terminal colors.
CI integration
GitHub Actions
name: Evals
on: [pull_request]
permissions:
contents: read
jobs:
evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- uses: superintelligenceco/evalkit@v0
id: evals
with:
suite: evals.yaml
python-version: "3.12"
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
- run: echo "Pass rate ${{ steps.evals.outputs.pass-rate }}"
The release workflow moves a major version tag, such as v0, to each new release, so
superintelligenceco/evalkit@v0 tracks the latest 0.x release. Pin the action to a full commit
SHA in production workflows. The action writes results.json,
junit.xml, summary.md, and badge.json to output-dir (default evalkit-results) and appends
the Markdown summary to the job summary.
| Input | Default | Description |
|---|---|---|
suite |
required | Path to the suite file. |
baseline |
Results JSON to compare against. | |
threshold |
0 |
Allowed drop against the baseline, as a fraction. |
strict |
false |
Fail when any single task regresses. |
fail-under |
Minimum pass rate. Overrides the suite setting. | |
output-dir |
evalkit-results |
Where the reports go. |
args |
Extra evalkit run arguments, such as --tag smoke. |
|
python-version |
Install this Python first. Leave empty to use the runner's python3. |
|
install |
the action's own copy | A pip requirement to install instead. |
job-summary |
true |
Write the Markdown summary to the job summary. |
Outputs: passed, total, pass-rate, ok, and results-file.
This repository runs its own examples through the action on every push. See
.github/workflows/evals.yml.
Baselines in CI
A simple pattern: upload results.json as an artifact on main, download the latest one in pull
request workflows, and pass it as baseline. The response cache makes the main-branch run cheap,
and you can cache .evalkit/ between jobs with actions/cache to skip repeat API calls.
Other CI systems
Run the CLI and publish the JUnit file with your system's test report feature. Download a release executable or install with pip:
pip install sic-evalkit
evalkit run evals.yaml --junit evalkit-junit.xml -o results.json
Badge
--badge badge.json writes a shields.io endpoint badge such as
{"schemaVersion": 1, "label": "evals", "message": "7/8 passed", "color": "yellow"}. Publish the
file anywhere public, such as GitHub Pages or a gist, and point
https://img.shields.io/endpoint?url=<url-of-badge.json> at it.
Results format
-o results.json writes everything evalkit knows about a run: the suite, target, seed, a summary
(pass rate, score, flaky and errored counts, tokens, cost, p50 and p95 latency, cached runs), and
every task with every run, output, grade, and reason. format: 1 versions the layout.
compare reads only the summary and each task's id, status, passed, and score.
Roadmap
These are planned, not built:
- Multi-turn conversation tasks and tool-call assertions for agents.
- More graders: semantic similarity with embeddings, ROUGE and BLEU, and a pairwise judge.
- Side-by-side comparison of several targets in one run.
- An HTML report with per-task diffs.
Suggestions are welcome in issues.
Contributing
Read CONTRIBUTING.md for setup, the project layout, and how to add a grader or a provider. Every test runs offline, so you can work on evalkit without an API key. Report security problems as described in SECURITY.md.
License
Release files for sic-evalkit 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sic_evalkit-0.2.0.tar.gz | 63.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sic_evalkit-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 114.3 kB
Release files / sic_evalkit-0.2.0.tar.gz
| Download URL | sic_evalkit-0.2.0.tar.gz |
|---|---|
| Size | 63.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
df13c0dccae4864b237eb1e6a920b0ff2634a61d02493c1ad2930b4be8941ad6
|
|
BLAKE2b-256 checksum How to use checksums |
2e58c2ae6c04342eb0c38de1df6860a9fae96b2c07825aeff3bef831c239ce92
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / sic_evalkit-0.2.0-py3-none-any.whl
| Download URL | sic_evalkit-0.2.0-py3-none-any.whl |
|---|---|
| Size | 51.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7acba605dd3a3892d3b171d0bef229a969b41ce8b0061e569096fce26cb5744c
|
|
BLAKE2b-256 checksum How to use checksums |
a5a765dd54ca27b494348f3ad2d07316da3a55918386d69540868959c95aad79
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|