Skip to main content

evaling

CI

A command-line tool for comparing prompt variants and models, easily.

Status: early development — alpha testers wanted. It works end to end and the test suite is thorough, but it's early: expect gaps where the tests don't reach, and expect the config format to shift a little before 1.0. If you try it, open an issue — rough edges, confusing output, and missing providers are all worth reporting.

Why

When you're iterating on a prompt, you want fast answers to questions like:

  • Which of these three phrasings performs best on my test cases?
  • Does the cheaper model handle this prompt as well as the expensive one?
  • Did my latest prompt tweak regress anything?

evaling runs your prompt variants against your chosen models over a set of test cases, scores the outputs, and shows you a comparison — all from the terminal.

Answering those questions with evidence, rather than by re-reading a handful of outputs, is usually what stands between a prompt change and shipping it. The aim is to make that comparison cheap enough that you run it every time.

What you can evaluate

Anything you can call. Anthropic and OpenAI directly; Ollama, vLLM, LM Studio, OpenRouter, Gemini and Gemma through the OpenAI-compatible endpoint; and anything else at all through the command provider, which pipes JSON to a script — so an agent, a RAG pipeline, or a local binary is as evaluable as a chat API.

Text, and not only text. Single prompts or scripted multi-turn conversations. Images, PDFs, audio, and video attach to cases and are checked against each model's declared capabilities before a request is sent.

Against whatever "correct" means for you. Exact match, substring, regex, and JSON-schema validation for the mechanical parts; your own Python function when the rule is real but not a pattern; an LLM judge with a rubric when the question is genuinely a matter of judgment — and a meta-eval to check that judge against labels you wrote yourself, before you trust it to gate anything.

At the size you actually have. A dozen cases inline in the config, a JSONL or CSV dataset in git, or hundreds of thousands streamed from your own warehouse or API through a case source — Python you write that evaling pages through. Memory is bounded by concurrency, not by case count, so a 500,000-case run costs no more up front than a ten-case one.

Including data you're not allowed to read. For the narrower case where the data is off-limits to humans, no-look mode — eyes-off evaluation, in privacy terms — keeps prompts, outputs, and attachments out of every artifact: you get the scores, the data stays where it was.

And it tells you whether the change helped. Compare two runs cell by cell, pin a baseline, and fail CI on a regression rather than on a hunch.

Install

Needs Python 3.10+ and uv:

uv tool install evaling
evaling --version

Or from a checkout, which is also the contributor setup:

git clone https://github.com/amro/evaling && cd evaling
uv tool install .

Working on evaling itself? Use uv sync and uv run evaling instead, which always reflects your working tree. See getting started.

Quick taste

evaling init      # scaffold a working example (offline, mock provider)
evaling run       # run it: progress bar, summary matrix, pass/fail gate
evaling show latest --failures
evaling compare <run-a> <run-b>
evaling export latest --format md

See getting started.

Documentation

Full docs live in docs/. New here? Start with the tutorial — install through CI gating, with runnable examples at every step.

  • Tutorial — the complete walkthrough.
  • Getting started — the short version.
  • CLI reference — commands, flags, run references, exit codes.
  • Configuration — the eval.yaml reference, settings layering, and environment variables.
  • Prompts — variants, Jinja2 templating, multi-turn messages, multimodal inputs, and case datasets.
  • Scoring — scorers, scorecards, LLM judges, and CI thresholds.
  • Providers — Anthropic, OpenAI, OpenAI-compatible (local models, Gemini, OpenRouter), the command provider, pricing, and adding your own.
  • Secrets — where API keys come from, and how they're kept out of git and out of your output.
  • Storage — the run directory format, resuming interrupted runs, response caching, and programmatic access.
  • Large datasets — case sources: streaming cases from your own API or warehouse instead of a file.
  • No-look evals — evaluating data nobody is permitted to read.
  • CI recipes — gating, baselines, cost ceilings, report artifacts.
  • Evaluating judges — calibrate autoraters against human labels (meta-evals).
  • MCP server — drive evaling from an agent for hands-off prompt iteration.
  • Claude Code plugin — install the MCP server, a command, and the eval workflow as a plugin.
  • Python API — use evaling as a library.
  • Troubleshooting — symptoms, causes, fixes.
  • Architecture — how it's built, and why.

Development

Requires uv.

uv sync              # create the venv and install dependencies
uv run evaling       # run the CLI
uv run pytest        # run the test suite
uv run ruff check .  # lint
uv run ruff format . # format

Changes are documented in the CHANGELOG.

Seven complete sample evals live in examples/ — text and multimodal, single- and multi-turn, a RAG pipeline behind the command provider, and a no-look run over a paging source. The test suite runs them all end to end on every commit, so they work with the current version.

See CONTRIBUTING.md before opening a pull request.

How this was built

evaling was written with heavy use of AI coding tools — most of it with Claude Code, working from requirements and design decisions recorded in REQUIREMENTS.md. Every change is reviewed before it lands, carries tests, and has to pass the same CI as anything else: lint, formatting, a dependency audit, and the full suite on Linux across Python 3.10–3.13, plus macOS and Windows. The docs and the worked examples are exercised by that suite too, so they cannot quietly drift from the code.

Worth saying plainly rather than leaving you to infer it from the commit log.

License

MIT; © 2026 Amro Mousa

Release files for evaling 0.2.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for evaling 0.2.4
File Size Uploaded
evaling-0.2.4.tar.gz 512.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for evaling 0.2.4
File Interpreter ABI Platform
evaling-0.2.4-py3-none-any.whl Python 3 none any Details

Total release size: 660.1 kB

Release files / evaling-0.2.4.tar.gz

Download URL evaling-0.2.4.tar.gz
Size 512.8 kB
Tags Source
SHA-256 checksum
How to use checksums
369c5a7008fc3bcadc1240b88adc88a6c9d2b9f69112760752153a1ebaac189e
BLAKE2b-256 checksum
How to use checksums
11dcf46b79ec063cca530a62e2131347f1d3e7944cce3aea737269e21ffc6692
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 17, 2026.

Transparency log

Release files / evaling-0.2.4-py3-none-any.whl

Download URL evaling-0.2.4-py3-none-any.whl
Size 147.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c35e87b53f0c84be0de1264855b10894b29a63296e0fd192e33e2abbb7704d2a
BLAKE2b-256 checksum
How to use checksums
ee738a8b832b02233f4ca037a83c232b9c3da8d03bbe6f0e2b3c51d9e51a74bb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 17, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

This release

0.2.4 This release

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page