digline
Regression testing for LLM applications — with the baseline in your repository, not on someone's server.
Your prompt worked on Tuesday. On Thursday it works a little less — not enough to break, enough for a user to notice in two weeks. No ordinary test catches it, because there is no correct output to compare against, only a better or a worse one.
digline gives you an approved reference — the baseline — and on every change
tells you whether you are below it: which case, which check, by how much. The
baseline is a JSON file in your repository, so it goes through code review and
it rolls back with git. No server, no account, no network call you have not
configured yourself.
$ digline compare --suite suite.py --run latest
2 checks got worse compared with the reference. Every case could be judged. No case is suspended. The configuration is the same as the reference.
how-do-i-return · llm_rubric · Score fell from 1.000000 to 0.700000.
how-do-i-return · contains · Went from passing to failing (1.000000 → 0.000000).
Why digline
Most evaluation tools tell you whether an output is below a threshold. digline also tells you whether it is worse than it was — the drift from 0.91 to 0.78 that trips no threshold and is the first thing a user feels.
The suite is Python, not YAML: a judge is an object, a target is a function, and what may leave a perimeter is declared in code — none of which a configuration file expresses without reinventing a language. Built for teams shipping LLM features for someone else, who have to show a customer what was tested, when, under which commit, and who approved it.
Quickstart
pip install digline digline-anthropic # or: uv add digline digline-anthropic
suite.py — complete and runnable, no API key:
"""suite.py — complete and runnable: no API key, nothing else to install."""
from digline.core import Contains, CostBudget, JudgeReply, LlmRubric
from digline.run import Case, Response, Suite
ANSWERS = {
"where-is-my-order": "Order 4821 ships Thursday. — Northwind Support",
"how-do-i-return": "Any item, within 30 days, unused. — Northwind Support",
}
def judge(prompt: str) -> JudgeReply:
"""Your judge. digline composes `prompt` from the rubric, the question and
the answer; it wants a score in [0, 1] and a reason back."""
signed = "Northwind Support" in prompt
concise = len(prompt.split()) <= 60
return JudgeReply(
score=0.4 + 0.3 * signed + 0.3 * concise,
reason=f"signed={signed}, concise={concise}",
)
def target(case: Case) -> Response:
"""Your application, called once per case. Canned here so this runs as is."""
text = ANSWERS[case.id]
return Response(output=text, cost_usd=0.004 + 0.001 * len(text) / 100)
suite = Suite(
tenant="northwind",
environment="staging",
name="support",
assertions=[
Contains(needle="Northwind Support"),
LlmRubric(
rubric="Does the reply answer the question in at most three sentences?",
judge=judge,
threshold=0.7,
tolerance=0.05,
),
CostBudget(max_usd=0.02, tolerance=0.05),
],
cases=[Case(id="where-is-my-order"), Case(id="how-do-i-return")],
)
$ digline run --suite suite.py
2026-08-26T15-44-09-282929-00-00-e7421ec503ccefe8
$ digline promote --suite suite.py --run latest
support baseline set to 2026-08-26T15-44-09-282929-00-00-e7421ec503ccefe8
Now change the prompt, the model, an answer — anything — and ask again:
$ digline run --suite suite.py
2026-08-26T15-44-09-492722-00-00-e7421ec503ccefe8
$ digline compare --suite suite.py --run latest
2 checks got worse compared with the reference. Every case could be judged. No case is suspended. The configuration is the same as the reference.
how-do-i-return · llm_rubric · Score fell from 1.000000 to 0.700000.
how-do-i-return · contains · Went from passing to failing (1.000000 → 0.000000).
$ echo $?
1
The exit code is the answer: 0 fine, 1 got worse, 2 could not be judged.
Everything lands in .digline/<tenant>/ — baselines/ committed, runs/
git-ignored through a .gitignore digline writes for you.
What it checks
Per case — pure functions (inputs) -> Verdict, no I/O, callable on their own:
| Assertion | Use it when |
|---|---|
Equals, Contains, NotContains, Affix, Regex |
the output must, or must not, contain something specific |
IsJson, JsonSchema |
the output is structured |
Length |
answers are growing, or must fit a channel |
Levenshtein |
"close enough" to Case.expected, graded rather than binary |
LlmRubric |
the criterion is a judgement — is it polite, does it stay on policy |
Faithfulness |
RAG: is the answer supported by the retrieved context |
FromAutoevals |
you already have an autoevals scorer and want it under a baseline |
PiiAbsent |
the output reaches a person — IBAN, codice fiscale, partita IVA, email, phone, checksum-verified where one exists |
CostBudget, LatencyBudget |
always. Graded, so a cost creeping up within budget is still visible |
Repeated |
the judge oscillates: grade the same output n times and fold the votes |
Per run — one verdict on the whole suite, the kind that goes in a contract:
| Aggregate | Use it when |
|---|---|
Precision |
false positives are what your users see |
Recall |
what is missed is what your users miss |
Accuracy, F1 |
you need a single number for both |
Every assertion carries a threshold that can fail — there is no default that
passes vacuously, and Contains("") is a ValueError when the suite loads
rather than a green run — and a tolerance below which a difference from the
baseline is noise. Where a number is really "k out of n", write it as one:
min_agreement="2/3", and a float no k/n can produce is refused at
construction.
One card each — parameters, typical values, what to watch out for — in
docs/metrics.md. Custom assertion? Subclass
AssertionBase, or RunAssertionBase for an aggregate: docs/api.md.
How it thinks
- The judge is yours. digline never calls a model API: you inject a function, and in your tests you inject a deterministic one.
- Three states, not two —
pass,fail,error. An error is neither green nor a regression: it means could not judge, and a run containing one cannot become the baseline. - Two kinds of noise, two answers.
Suite.samplesasks the target more than once — the same input answered differently.Repeatedgrades the same output more than once — the judge changing its mind.min_agreementbecomes mandatory as soon as you sample. - Set the threshold where the system measurably is, not where you want it: the gate protects against getting worse, and raising the bar is a visible change in a pull request.
- Promote the median of several runs, not the first green one —
digline viewis the table you pick it from. Cases diagnose, aggregates gate.
Worked through with real numbers in docs/guide.md; the
reasoning behind every fixed decision is in docs/adr/.
Commands
| Command | |
|---|---|
digline run |
execute the suite, write the run, print its key |
digline compare |
headline plus the lines that got worse; --json, --json full for CI |
digline promote |
make a run the baseline — refused if the tenant differs, the configuration changed, or any check errored |
digline report |
self-contained HTML for readers who do not read code; --locale mandatory, --redacted keeps the verdicts and drops the payload |
digline list |
stored runs, newest first, baseline marked |
digline view |
local browser UI — docs/view.md |
digline migrate |
bring stored runs forward across schema versions — docs/migrate.md |
What digline is not
- Not an observability platform. Dashboards over production traces are a served market. What is designed and not yet built is narrower: evaluating production responses inside your perimeter, and turning a failure into a committed test case.
- Not a red-teaming tool. digline generates no attacks. Once one is found,
it becomes a
Case, and the suite makes sure it never works again. - Not YAML. Cases are data and may come from files; the suite is Python.
Status
0.1.0, alpha. The offline cycle — write the suite, run, promote, compare,
report — is complete, covered by tests, and used daily on a real project. The
production store, the bridge from production failures back to committed cases,
and the reactive side are designed in
ADR 0002 and not
written yet.
Python 3.12+. One runtime dependency: jsonschema.
Docs
docs/guide.md— how to reason with digline, in eight chapters and the order the problems arrive: baseline, judge noise, sampling, tolerance, threshold, which run to promote, what to gate on, what to maintaindocs/metrics.md— a card per assertion and aggregate: when to reach for it, what it produces, what it will do to you if you are not lookingdocs/api.md— what is imported from where, every assertion and its parameters, custom assertions, and the complete example inexamples/quickstart/, which a test runs on every builddocs/view.md·docs/migrate.md— the two commands with a surface of their owndocs/adr/— the architectural decisions, numbered, with the reasoning
License
Apache-2.0.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file digline-0.1.1.tar.gz.
File metadata
- Download URL: digline-0.1.1.tar.gz
- Upload date:
- Size: 243.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5a5c773d6b2f7fb32ec9df1687af15494b4a928e4e656c38fe6abe583585887d
|
|
| MD5 |
299e58ef08711ea79c47f9974bb2c6dc
|
|
| BLAKE2b-256 |
895907f990b1b5e10e70a8e35379766584e00d18c6c941c1a99817aec26ec182
|
Provenance
The following attestation bundles were made for digline-0.1.1.tar.gz:
Publisher:
publish.yml on digline/digline
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
digline-0.1.1.tar.gz -
Subject digest:
5a5c773d6b2f7fb32ec9df1687af15494b4a928e4e656c38fe6abe583585887d - Sigstore transparency entry: 2615320571
- Sigstore integration time:
-
Permalink:
digline/digline@b5870a655c776d75f8be518796bad8acf02fb08a -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/digline
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b5870a655c776d75f8be518796bad8acf02fb08a -
Trigger Event:
push
-
Statement type:
File details
Details for the file digline-0.1.1-py3-none-any.whl.
File metadata
- Download URL: digline-0.1.1-py3-none-any.whl
- Upload date:
- Size: 130.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2add81fc790a2ba3e3a878f518c45a4973e998defc66d5b402d619b22791b3a2
|
|
| MD5 |
511c7702ba7474f5342654705dd52d0f
|
|
| BLAKE2b-256 |
3a712c43f11dcb93412ef76cc17fd11f5d7b8e8adad91aaefce0b132b0030a8c
|
Provenance
The following attestation bundles were made for digline-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on digline/digline
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
digline-0.1.1-py3-none-any.whl -
Subject digest:
2add81fc790a2ba3e3a878f518c45a4973e998defc66d5b402d619b22791b3a2 - Sigstore transparency entry: 2615321037
- Sigstore integration time:
-
Permalink:
digline/digline@b5870a655c776d75f8be518796bad8acf02fb08a -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/digline
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b5870a655c776d75f8be518796bad8acf02fb08a -
Trigger Event:
push
-
Statement type: