Skip to main content

evalix

A small eval harness that tells you which cases your prompt change broke.

  score      0.136 → 0.136  (+0.000)   over 22 shared cases
    fixed  t11
    BROKE  t14

A flat score. One case fixed, one case broken. Without the second line you would have called that "no change" and moved on.

That is the whole pitch. The score tells you whether; the diff tells you which, and only the second one teaches you anything.

The loop

fixed case set → change one thing → measure → keep or revert

The usual way people practise prompt engineering is to tweak a prompt, look at one output, decide it's better, and move on. That's not a skill, it's a feeling. Everything here is built around the loop above.

Install

pip install evalix        # or: uv add evalix

The default runner talks to Claude through capix, so ANTHROPIC_API_KEY in your environment or a .env file is all the setup there is. Any other provider is a three-line function — see Runners.

Use it

A case file is JSONL, one case per line:

{"id": "t01", "tag": "easy", "input": "I was charged twice.", "expected": "billing"}
{"id": "t02", "tag": "edge", "input": "Export CSV throws a TypeError.", "expected": "bug"}
evalix run --cases cases.jsonl --prompt prompts/v1.txt --scorer exact
  ✓ t01        1.00  want='billing' got='billing'
  ✗ t02        0.00  want='bug' got='how_to'

  score      0.727   (16/22 perfect)
  tokens     1534 in / 212 out   ≈ $0.0026

  0.636 → 0.727  (+0.091)   over 22 shared cases
    fixed  t09, t11, t14
    BROKE  t20

  worst 5:
    t20 (0.00) want='how_to' got='feature_request'
        → feature_request

Runs are saved to runs/ and the next run of the same case file is diffed against them automatically. To compare any two runs — v1 against v5, with three experiments in between:

evalix compare v1-lazy v5-spec

Each argument is a run file or any substring of one. Reruns of the same experiment resolve to the newest; a substring that matches different runs (v1 also matching v10) is an error rather than a guess. It prints the metadata side by side, flags every field that differs, warns when more than one axis moved at once, then gives the per-case changes, a per-tag breakdown, and the outputs behind each regression.

Spend nothing while you work

--dry-run sends nothing. It resolves the case file, the prompt and the scorer, prints what each call would contain, and estimates the bill:

evalix run --cases cases.jsonl --prompt prompts/v1.txt --dry-run
  22 calls · ~3211 input tokens
  worst case 44000 output tokens (every case hitting max-tokens)
  estimated  $0.0032 … $0.2232   (input only … input + max output)

  dry run — nothing was sent, no run file written

With --scorer judge the estimate covers the run only; the judge's calls come on top. Use it whenever you edit a prompt or a case file. A mangled {input} placeholder or a prompt that grew tenfold shows up here for free.

Scorers

A scorer turns one output into a number. It is the only task-specific part of the harness, which is why it is the pluggable one — and the part most likely to be silently wrong, which is why the built-ins are tested to the letter.

--scorer expected is scores
exact a string 1 if equal after normalising case, whitespace, trailing punctuation
contains a string 1 if present as a substring
not_contains a string 1 if absent — the injection canary
regex a pattern 1 if it matches
json_parse (unused) 1 if the output parsed as JSON at all
json_fields an object the fraction of fields that match; unparseable scores 0
judge optional reference answer a second model call against --rubric, 1–5 mapped onto 0…1
none (unused) nothing — records the output for manual reading

Any key you add to a case is passed through untouched, so a custom scorer can read whatever its task needs:

# score.py — canary absent AND the real job still done
def score(output, case):
    if case["expected"].lower() in output.lower():
        return 0.0, "LEAKED"
    anchors = case.get("must_contain", [])
    if not anchors:
        return 1.0, "clean"
    hit = [a for a in anchors if a.lower() in output.lower()]
    return len(hit) / len(anchors), "clean" if len(hit) == len(anchors) else "off-task"
evalix run --cases cases.jsonl --scorer custom --scorer-file score.py

A scorer that raises stops the run rather than scoring the case zero, and cases that haven't started yet are cancelled, not paid for. A broken scorer otherwise reports a clean 0.000 that looks exactly like a failing prompt. --keep-going opts out.

Take a third argument to get a Context — with it, your scorer can call a model itself, which is all a judge is:

from evalix import Message, Request

def score(output, case, ctx):
    verdict = ctx.runner(Request(messages=[Message("user", f"Grade this: {output}")]))
    ...

Calls made through ctx.runner are counted: their tokens show on their own scorer line and their cost is included in the total. If one of them fails with a transport error, the case is recorded as unscored and the run continues.

Python API

The CLI is a thin wrapper over this, so nothing is terminal-only:

from evalix import run

report = run(
    cases="cases.jsonl",
    prompt="prompts/v2.txt",
    scorer="json_fields",
    model="claude-haiku-4-5",
)

report.score          # 0.727
report.results        # list[Result] — score, note, output, latency, tokens
report.cost_usd       # includes any model calls the scorer made
print(report.render())

if report.diff:       # None on the first run of a case file
    report.diff.broke # ['t20']

compare is the same diff for any two saved runs:

from evalix import compare

result = compare("v1-lazy", "v5-spec")
result.diff.broke     # ['t14']
result.old.meta       # everything recorded about the baseline run
print(result.render())

Runners

A runner is any callable taking a Request and returning a Response. Nothing in the core imports capix, so another provider is an argument, not a fork:

from evalix import run, Response

def my_runner(request):
    reply = my_client.complete(system=request.system, prompt=request.text)
    return Response(text=reply, input_tokens=..., output_tokens=...)

run(cases="cases.jsonl", runner=my_runner, model="whatever-you-call-it")

It is also how the whole pipeline gets tested offline — every test in this repo runs without an API key, because a fake runner is three lines.

v1 is single-turn. Request.messages is a list so that tools and extra turns can arrive as a field rather than a new major version, but nothing sends more than one message today.

Flags

evalix run
  --cases FILE           JSONL, one case per line
  --prompt FILE          the prompt under test
  --scorer NAME          exact | contains | not_contains | regex | json_parse
                         | json_fields | judge | none | custom
  --scorer-file FILE     with --scorer custom: a .py defining score()
  --rubric FILE          with --scorer judge
  --judge-model ID       default claude-opus-5
  -m / --model ID        model under test
  -e / --effort LEVEL    low|medium|high|xhigh|max (thinking-capable models)
  -t / --max-tokens N    default 2000
  --placement system|user  where the prompt file goes
  --repeat N             run each case N times — consistency check
  --only-tag TAG         slice to one tag
  --limit N              first N cases only
  --label NAME           name this run in the diff output
  --workers N            parallel requests (default 8)
  --runs-dir DIR         default <project root>/runs, or $EVALIX_RUNS
  --keep-going           score a case zero when the scorer raises
  --dry-run              print what would be sent, send nothing
  --quiet / --show N     less per-case noise / how many failures to print

evalix compare OLD NEW [--all] [--show N] [--runs-dir DIR]

Habits worth stealing

  • Change one thing per run. Two changes and a flat score tells you nothing. compare warns when more than one axis moved.
  • Keep every prompt version. v1-lazy.txt, v2-spec.txt — they're the log of what you learned, and you'll want to revert.
  • Read the failures, not the score. The score tells you whether, the outputs tell you why.
  • Suspect your labels. When a case won't budge, check whether your expected answer is actually right. Sometimes the model is and you aren't.
  • Check consistency before celebrating. --repeat 3. A case that flips between runs isn't solved, it's lucky.

Development

uv sync
uv run pytest -q
uv run ruff check

The whole suite is offline — no key, no tokens, nothing spent.

Licence

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

evalix-0.1.0.tar.gz (36.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

evalix-0.1.0-py3-none-any.whl (34.6 kB view details)

Uploaded Python 3

File details

Details for the file evalix-0.1.0.tar.gz.

File metadata

  • Download URL: evalix-0.1.0.tar.gz
  • Upload date:
  • Size: 36.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for evalix-0.1.0.tar.gz
Algorithm Hash digest
SHA256 1fc93795fb3d43f2612eae7ba3542d38edf437bb00d54f10c352cae9d760f27c
MD5 d1c202d0540c7ae39ba245b6c8365009
BLAKE2b-256 b1c7a5c793999ec99fbfe405ad9b4e118e3b9ab9221b733f7c2388bafe090d72

See more details on using hashes here.

File details

Details for the file evalix-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: evalix-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 34.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for evalix-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b840e538a6ab62db371c366efc7df26565e6cf93b35e7ced6b465e47ff15652b
MD5 7c676419f7992a4f0de932db579301e4
BLAKE2b-256 62765f77265220a5fc668aee6bb0979e033fe540eaacf3aa46914219fc3ce50b

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page