Skip to main content

digline

Regression testing for LLM applications — with the baseline in your repository, not on someone's server.

PyPI Python 3.12+ License: Apache-2.0 CI

Your prompt worked on Tuesday. On Thursday it works a little less — not enough to break, enough for a user to notice in two weeks. No ordinary test catches it, because there is no correct output to compare against, only a better or a worse one.

digline gives you an approved reference — the baseline — and on every change tells you whether you are below it: which case, which check, by how much. The baseline is a JSON file in your repository, so it goes through code review and it rolls back with git. No server, no account, no network call you have not configured yourself.

$ digline compare --suite suite.py --run latest
2 checks got worse compared with the reference. Every case could be judged. No case is suspended. The suite is unchanged from the reference.

how-do-i-return · llm_rubric · Score fell from 1.000000 to 0.700000.
how-do-i-return · contains · Went from passing to failing (1.000000 → 0.000000).

Why digline

Most evaluation tools tell you whether an output is below a threshold. digline also tells you whether it is worse than it was — the drift from 0.91 to 0.78 that trips no threshold and is the first thing a user feels.

The suite is Python, not YAML: a judge is an object, a target is a function, and what may leave a perimeter is declared in code — none of which a configuration file expresses without reinventing a language. Where a suite is plain data it can be TOML instead, and the two forms build the same objects — what TOML cannot express, it refuses by name rather than half-supporting. Built for teams shipping LLM features for someone else, who have to show a customer what was tested, when, under which commit, and who approved it.

Wondering how digline differs from promptfoo, DeepEval, or observability platforms? See How digline compares.

Quickstart

With uv (recommended):

uv init && uv add digline

or with pip in an existing environment: pip install digline.

Requires Python 3.12+, which uv fetches for you if you do not have it. On the pip path an older interpreter says ERROR: No matching distribution found with from versions: none, which does not say why — that is what it means.

suite.py — complete and runnable, no API key:

"""suite.py — complete and runnable: no API key, nothing else to install."""

from digline.core import Contains, CostBudget, JudgeReply, LlmRubric
from digline.run import Case, Response, Suite

ANSWERS = {
    "where-is-my-order": "Order 4821 ships Thursday. — Northwind Support",
    "how-do-i-return": "Any item, within 30 days, unused. — Northwind Support",
}


def judge(prompt: str) -> JudgeReply:
    """Your judge. digline composes `prompt` from the rubric, the question and
    the answer; it wants a score in [0, 1] and a reason back."""
    signed = "Northwind Support" in prompt
    concise = len(prompt.split()) <= 60
    return JudgeReply(
        score=0.4 + 0.3 * signed + 0.3 * concise,
        reason=f"signed={signed}, concise={concise}",
    )


def target(case: Case) -> Response:
    """Your application, called once per case. Canned here so this runs as is."""
    text = ANSWERS[case.id]
    return Response(output=text, cost_usd=0.004 + 0.001 * len(text) / 100)


suite = Suite(
    tenant="northwind",
    environment="staging",
    name="support",
    assertions=[
        Contains(needle="Northwind Support"),
        LlmRubric(
            rubric="Does the reply answer the question in at most three sentences?",
            judge=judge,
            threshold=0.7,
            tolerance=0.05,
        ),
        CostBudget(max_usd=0.02, tolerance=0.05),
    ],
    cases=[Case(id="where-is-my-order"), Case(id="how-do-i-return")],
)

When your judge is a real model, add a provider plugin: uv add digline-anthropic (or pip install digline-anthropic), likewise digline-openai and digline-bedrock.

$ digline run --suite suite.py
2026-08-26T15-44-09-282929-00-00-e7421ec503ccefe8

$ digline promote --suite suite.py --run latest
support baseline set to 2026-08-26T15-44-09-282929-00-00-e7421ec503ccefe8

Now make it worse — delete — Northwind Support from the second answer — and ask again:

$ digline run --suite suite.py
2026-08-26T15-44-09-492722-00-00-e7421ec503ccefe8

$ digline compare --suite suite.py --run latest
2 checks got worse compared with the reference. Every case could be judged. No case is suspended. The suite is unchanged from the reference.

how-do-i-return · llm_rubric · Score fell from 1.000000 to 0.700000.
how-do-i-return · contains · Went from passing to failing (1.000000 → 0.000000).

$ echo $?
1

The exit code is the answer: 0 fine, 1 got worse, 2 could not be judged. Everything lands in .digline/<tenant>/ — baselines/ committed, runs/ git-ignored through a .gitignore digline writes for you.

What it checks

Per case — pure functions (inputs) -> Verdict, no I/O, callable on their own:

Assertion Use it when
Equals, Contains, NotContains, Affix, Regex the output must, or must not, contain something specific
IsJson, JsonSchema the output is structured
Length answers are growing, or must fit a channel
Levenshtein "close enough" to Case.expected, graded rather than binary
LlmRubric the criterion is a judgement — is it polite, does it stay on policy
Faithfulness RAG: is the answer supported by the retrieved context
FromAutoevals you already have an autoevals scorer and want it under a baseline
PiiAbsent the output reaches a person — IBAN, codice fiscale, partita IVA, email, phone, checksum-verified where one exists
CostBudget, LatencyBudget always. Graded, so a cost creeping up within budget is still visible
Repeated the judge oscillates: grade the same output n times and fold the votes

Per run — one verdict on the whole suite, the kind that goes in a contract:

Aggregate Use it when
Precision false positives are what your users see
Recall what is missed is what your users miss
Accuracy, F1 you need a single number for both

Every assertion carries a threshold that can fail — there is no default that passes vacuously, and Contains("") is a ValueError when the suite loads rather than a green run — and a tolerance below which a difference from the baseline is noise. Where a number is really "k out of n", write it as one: min_agreement="2/3", and a float no k/n can produce is refused at construction.

One card each — parameters, typical values, what to watch out for — in docs/metrics.md. Custom assertion? Subclass AssertionBase, or RunAssertionBase for an aggregate: docs/api.md.

How it thinks

  • The judge is yours. digline never calls a model API: you inject a function, and in your tests you inject a deterministic one.
  • Three states, not two — pass, fail, error. An error is neither green nor a regression: it means could not judge, and a run containing one cannot become the baseline.
  • Two kinds of noise, two answers. Suite.samples asks the target more than once — the same input answered differently. Repeated grades the same output more than once — the judge changing its mind. min_agreement becomes mandatory as soon as you sample.
  • A tolerance is declared; a noise floor is measured. A sampled run records the interval its own samples spanned, and a drop that stays inside the baseline's interval is reported as unchanged rather than as a regression — a tool that cries wolf on its own measurement error teaches people to promote past it. It never rescues a flip, and it never invents an interval it does not have.
  • Set the threshold where the system measurably is, not where you want it: the gate protects against getting worse, and raising the bar is a visible change in a pull request.
  • Promote the median of several runs, not the first green one — digline view is the table you pick it from. Cases diagnose, aggregates gate.

Worked through with real numbers in docs/guide.md; the reasoning behind every fixed decision is in docs/adr/.

Commands

Command
digline run execute the suite, write the run, print its key
digline compare headline plus the lines that got worse; --json, --json full for CI
digline diff what differs between two runs, neither of them a baseline — for "should I switch?" rather than "did it get worse?". Always exits 0: it is a report, not a verdict — docs/diff.md
digline promote make a run the baseline — refused if the tenant differs, the configuration changed, or any check errored
digline report self-contained HTML for readers who do not read code; --locale mandatory, --redacted keeps the verdicts and drops the payload. With no baseline yet it renders the run on its own and says so, so the first run is readable before anything is promoted
digline explain the same facts read back at length, in prose: what ran, what moved, against which measured interval, what differed underneath. Compares when there is a baseline and reads the run alone when there is not. --json emits the fact list the prose is rendered from. It states, and never advises — docs/explain.md
digline list stored runs, newest first, baseline marked
digline view local browser UI — docs/view.md
digline migrate bring stored runs forward across schema versions — docs/migrate.md

Examples

Eight projects in examples/, each answering a question somebody actually arrives with. Every one runs with no API key, carries its committed report.html, and is a standalone project: copy the directory anywhere and uv sync works.

What digline is not

  • Not an observability platform. Dashboards over production traces are a served market. What is designed and not yet built is narrower: evaluating production responses inside your perimeter, and turning a failure into a committed test case.
  • Not a red-teaming tool. digline generates no attacks. Once one is found, it becomes a Case, and the suite makes sure it never works again.
  • Not YAML. The suite is Python — or, within declared limits, TOML: cases were always data, and now the rules can be too. A judge with rules of its own, a computed request body and a custom assertion stay Python, and the loader names the wall you hit rather than half-supporting it.
  • Not a funnel. Two commitments, by design and for good: no hosted service that receives your payloads, and no data collection. The baseline lives in your repo; the runs happen on your machines. If digline ever grows paid features, they will run inside your perimeter too.

Status

0.7.0, pre-1.0. The offline cycle — write the suite, run, promote, compare, report — is complete, covered by tests, and used daily on a real project, and since 0.5.0 the suite may be written as data as well as in Python. The API may still change before 1.0; the baseline format is versioned and migrates. The production store, the bridge from production failures back to committed cases, and the reactive side are designed in ADR 0002 and not written yet.

Python 3.12+. One runtime dependency: jsonschema.

Docs

  • docs/guide.md — how to reason with digline, in eight chapters and the order the problems arrive: baseline, judge noise, sampling, tolerance, threshold, which run to promote, what to gate on, what to maintain
  • docs/metrics.md — a card per assertion and aggregate: when to reach for it, what it produces, what it will do to you if you are not looking
  • docs/api.md — what is imported from where, every assertion and its parameters, custom assertions, and the complete example in examples/quickstart/, which a test runs on every build
  • docs/declarative.md — the suite as data: the TOML format, key by key, what it deliberately cannot say, and how to move a suite between the two forms without losing its baseline
  • docs/view.md · docs/migrate.md — the two commands with a surface of their own
  • AGENTS.md — how a coding agent should operate digline in your repo
  • docs/adr/ — the architectural decisions, numbered, with the reasoning

License

Apache-2.0.

Release files for digline 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for digline 0.7.0
File Size Uploaded
digline-0.7.0.tar.gz 821.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for digline 0.7.0
File Interpreter ABI Platform
digline-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / digline-0.7.0.tar.gz

Download URL digline-0.7.0.tar.gz
Size 821.8 kB
Tags Source
SHA-256 checksum
How to use checksums
16a27a1809d1a715d04ef0daa28a417cf98b5c5c5707525b29e707272e1112b4
BLAKE2b-256 checksum
How to use checksums
ba415b12ad89e0048957d8a7476d23ed8ceec2326c2c32763da1e1aa55a1ba93
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release files / digline-0.7.0-py3-none-any.whl

Download URL digline-0.7.0-py3-none-any.whl
Size 240.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e5725bac00b5760c10e9ed1e8d49656de703cb051b8ce05d49f8d6e7a91aa082
BLAKE2b-256 checksum
How to use checksums
3ca8d31a80643199c307fdb500ca81999cb6e6ee7abdada98c529406ea948df0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release history Release notifications | RSS feed

0.20.0

2 release files

0.19.2

2 release files

0.19.1

2 release files

0.19.0

2 release files

0.18.0

2 release files

0.17.1

2 release files

0.17.0

2 release files

0.16.0

2 release files

0.15.3

2 release files

0.15.2

2 release files

0.15.1

2 release files

0.15.0

2 release files

0.14.1

2 release files

0.14.0

2 release files

0.13.3

2 release files

0.13.2

2 release files

0.13.1

2 release files

0.13.0

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.2

2 release files

0.7.1

2 release files

This release

0.7.0 This release

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page