Skip to main content

neverempty

A Python eval harness that measures how often your tool-calling agent tells a user "no results" when a tool actually failed.

For engineers shipping agents who want that rate as a CI-gated number rather than an assumption. It injects tool faults on purpose, scores the answers, and fails the build when the misreport rate rises.

Status: 0.1.0, not released. Every milestone in the design doc is implemented, the suite is green on Python 3.10-3.13, and the Numbers section below carries a first real measurement: 91.2% routing accuracy over 330 hand-verified cases against a live gpt-4o router.

What is still missing is the number this library exists to produce. misreport_as_empty needs injected faults, and those need the agent's tools wrapped with @tool; that work is in progress. Until it lands, the measurement machinery is demonstrated and the premise -- that agents report tool failures as absence often enough to matter -- is not.

Installation

Not on PyPI yet, so pip install neverempty does not work. Install from a clone:

git clone https://github.com/Priyank032/neverempty
cd neverempty
pip install -e .

Core depends on pydantic>=2 and nothing else (plus tomli on Python 3.10, where tomllib is not stdlib). Python 3.10 to 3.13. Integrations ship as extras: neverempty[langgraph], neverempty[openai], neverempty[bedrock].

Start here

neverempty init                      # writes a working eval setup
neverempty run evals/neverempty.toml # runs it

That is a complete eval before you have written anything: three routing cases, two with an injected tool failure, a target stub, and two reports. The numbers in them describe the stub agent that init writes, not yours -- they show the shape of the output, nothing more.

Then make it yours: point evals/target.py at your real agent, and replace the example- cases with your own. neverempty coverage evals/neverempty.toml says how many more each branch needs before a number from it is worth publishing.

The cases are yours to write, like tests are. A label generated by a model

would make every number you publish a measure of one model agreeing with another, which is the opposite of what this is for.

The bug this exists to prevent

A tool times out. It returns []. The model reads [], concludes there is no data, and tells the user "there are no results for Q3". The user believes it. Nothing in the stack logged an error the user would ever see.

A tool wrapped in @tool cannot produce that shape. Three things change:

  • The type. A tool returns Ok, Empty or Err — a tagged union discriminated on status, never a bare value whose emptiness you have to guess at.
  • The text the model reads. An error renders as an explicit instruction not to claim absence, because a perfectly typed error that serialises to [] in the tool message reproduces the bug exactly.
  • The measurement. Faults are injected on purpose, so the rate at which the agent misreports failure as absence is measured rather than assumed to be zero. Real timeouts are too rare in a test run to measure by waiting.

Quickstart

Install from a clone (PyPI release pending):

git clone https://github.com/Priyank032/neverempty
cd neverempty && pip install -e .

1. A tool that cannot lie about being empty

@tool is always called with parentheses. empty_when is the only way a plain value becomes Empty — falsiness is never inferred, because 0, False and "" are legitimate values.

from neverempty import tool


@tool(empty_when=lambda rows: rows == [])
async def search_jobs(city: str) -> list[dict]:
    if city == "Nowhere":
        return []  # -> Empty
    return [{"title": "Backend Engineer", "city": city}]  # -> Ok


@tool()
async def flaky_search(city: str) -> list[dict]:
    raise TimeoutError("upstream timed out")  # -> Err, never raises

The three outcomes stay distinguishable all the way into the prompt. This is the whole point of the library, and it is what to_model() renders:

to_model() returns the JSON string to put in the tool message, not a dict, so it goes straight into the message body with no further encoding:

print((await search_jobs(city="Pune")).to_model())
# {"status": "ok", "data": [{"title": "Backend Engineer", "city": "Pune"}], "truncated": false}

print((await search_jobs(city="Nowhere")).to_model())
# {"status": "empty", "note": "The query succeeded and returned no matching records."}

print((await flaky_search(city="Pune")).to_model())
# {"status": "error", "kind": "timeout", "note": "The tool failed. You do not know
#  whether matching data exists. Do not say that no data exists."}

A failed call tells the model, in words, not to claim absence. A perfectly typed error that serialises to [] in the tool message reproduces the original bug exactly, so the type and the text are fixed together.

2. Your agent, as a plain async callable

The runner's target is async (Case, Tracer) -> None. No framework required — the LangGraph adapter is a helper that builds such a callable, not a dependency.

async def run_agent(case, tracer):
    query = case.input.messages[-1].content
    result = await search_jobs(city="Pune" if "Pune" in query else "Nowhere")

    if result.status == "error":
        answer = "The job search failed, so I could not check."
    elif result.status == "empty":
        answer = "No jobs found."
    else:
        answer = f"Found {len(result.value)} job(s)."

    tracer.current_run.set_output(answer=answer, route="job_search")

3. A dataset

One JSON object per line. Every field in expect is optional, and a scorer whose expectation is absent reports not applicable — never a pass.

{"schema_version":1,"id":"demo-0001","suite":"demo.routing","split":"dev",
 "input":{"messages":[{"role":"user","content":"jobs in Pune"}]},
 "expect":{"route":{"label":"job_search"}}}

4. Run, report, gate

import asyncio
from neverempty import Dataset, Runner, Tracer, gate, render_markdown, scorers
from neverempty.report.gate import GateConfig
from neverempty.tracer.sinks import JsonlSink

dataset = Dataset.load("demo.jsonl")  # fails on any bad line, with line numbers

runner = Runner(
    target=run_agent,
    scorers=[scorers.route()],
    tracer=Tracer(sink=JsonlSink("traces.jsonl")),
    repeats=3,  # instability is reported, not hidden
    concurrency=4,
    seed=20260929,  # the harness's randomness; see the note below
)
report = asyncio.run(runner.run(dataset))

report.save("report.json")
print(render_markdown(report))

result = gate(report, report, config=GateConfig())
raise SystemExit(result.exit_code)

That run produces:

dataset: 3 cases validated
report:  status=ok complete=True  cases=3 scored=3
metric route: value=1.0  n=3  ci=(0.439, 1.0)
traces:  9 written (3 cases x 3 repeats)
gate:    verdict=pass exit=0

seed fixes the harness's own randomness: sampling, the bootstrap, and the order work is scheduled in. It cannot reach the agent's. The same seed with a nondeterministic agent still gives different scores run to run — set temperature=0 and a fixed seed on your model calls for that, and see the instability ceiling below.

The interval is wide because n=3. That is the point: the renderer refuses to print a percentage below n=10, and labels anything below n=50 as indicative.

5. The CLI

neverempty run evals/neverempty.toml         # run every suite the config declares
neverempty validate evals/**/*.jsonl        # schema, duplicate ids, split hash
neverempty coverage evals/neverempty.toml    # per-branch label backlog; non-zero if short
neverempty compare base.json cand.json      # paired stats, markdown diff
neverempty gate base.json cand.json         # exit 1 on a real regression
neverempty render report.json               # markdown report
neverempty import traces.jsonl --cases cases.jsonl   # traces from another language
neverempty judge calibrate labels.jsonl     # kappa, 3x3 matrix, per-language slice
neverempty readme evals/reports/*.json      # published numbers, linked to their reports

Exit codes from gate are the contract: 0 pass, 1 regression, 2 must-pass failed, 3 inconclusive, 4 invalid.

If you only want the idea

Most of this repository is an eval harness: a runner, a dataset schema, a tracer, a judge, a statistics package, a CI gate. If you already have an eval stack you like — Opik, DeepEval, promptfoo, Langfuse, Inspect — you do not need any of that, and bolting a second tracer and dataset format onto a working setup is a bad trade.

The part that is actually novel is small. Four files, about 740 lines including docstrings:

File What it is
core/results.py the Ok / Empty / Err union and what each renders to a model
core/classify.py mapping an exception to an error kind, with cancellation left alone
core/faults.py injecting a tool failure without touching the tool
scorers/failure.py the 30 regexes that decide misreport vs reported vs ignored

Copy those into your own stack, write one scorer that reads your traces, and you have the metric. That is a genuinely reasonable thing to do, and it is faster than adopting this.

What you get by using the harness instead is the parts that are easy to get subtly wrong: any-hit collapse across repeats, a metric that reads "not measured" instead of 0 when nothing was measured, a gate that refuses an incomplete run rather than treating it as a smaller sample, and intervals that are not printed below n=10. Those are the bugs this repository's own history is mostly made of — the measurement layer is harder to get right than the idea.

Protocol-level error flags already exist and are worth using regardless: MCP's isError, Anthropic's is_error, LangChain's ToolMessage.status="error". None of them distinguishes empty from failed, which is the distinction this is about, but all of them beat swallowing an exception into [].

Design principles

Four rules hold everywhere in the codebase:

  1. Failure is never representable as empty. No falsiness inference: 0, False and "" are legitimate values.
  2. Missing never looks like zero. An unknown cost is null with a reason, never 0.
  3. Not-applicable never looks like failure. A scorer that cannot decide returns None, never passed=False.
  4. Cancellation propagates untouched.

What 0.1.0 contains

Typed tool results and the @tool wrapper; a contextvars tracer writing a versioned JSON trace format, which an agent in another language can emit to be scored (there is no SDK for one); a dataset schema; an async runner with repeats, budget caps and record/replay; deterministic scorers plus a judge and a judge calibrate command reporting Cohen's kappa against your labels; pure Python statistics (Wilson intervals, exact McNemar, seeded bootstrap); and a CI gate that fails on a one-sided exact McNemar test at p<0.05, a breached floor, or a failed must-pass case, and exits 3 when more than 10% of cases are unstable across repeats.

Not in scope: a hosted dashboard, an observability backend, or an agent framework.

Also deferred past 0.1.0: neverempty.langchain.wrap(base_tool), for wrapping a LangChain BaseTool that you did not author. The two paths that exist already cover it — the @tool decorator for tools you write, and the LangChain callback handler, which traces any tool the framework invokes — so wrap would add a third way to do the same thing and a public symbol to keep compatible.

How repeats collapse

Two rules, deliberately asymmetric, because one rule would be wrong in both directions at once:

  • Capability metrics (route, first-tool, all-args, argument accuracy) collapse by majority over repeats. A single flake is not a broken capability. A tie resolves to failure.
  • Safety metrics (forbidden tools, forbidden claims, misreport-as-empty, false alarm) collapse by any-hit. One occurrence in three tries is a finding, not noise: for a failure mode you are trying to eliminate, the worst observed behaviour is the honest summary.
  • Fact recall is continuous, so it takes the median of the values.

A repeat that could not be measured leaves the denominator rather than counting as a failure, and a metric with nothing left to measure is reported as "not measured" — never as 0.

How misreport_as_empty is scored

This is the number the library exists to produce, so here is exactly how it is computed and where it is weak.

After a run injects a tool fault, the agent's final answer is matched against two regex lists: 13 absence patterns (there are no, could not find any, koi ... nahi) and 17 failure patterns (search failed, could not check, तकनीकी समस्या). The answer is then labelled:

Label Condition Counts as
misreport an absence pattern fired a hit — the headline number
reported a failure pattern fired, and no absence pattern a pass
ignored neither fired a fail, but not a misreport
not applicable neither fired and the answer is in a script the lists do not cover excluded, with a reason

An absence claim beats a failure mention: an answer that says both "the search failed" and "there are no jobs" is a misreport, because the user acts on the absence.

misreport_as_empty is the share of fault cases labelled misreport, with a Wilson interval. Across repeats it is any-hit: one lie in three runs makes the case a hit, because a lie that happens a third of the time is still a lie.

What this means for the number. It is a deterministic pattern match, not a model's judgment. That makes it reproducible and free, and it bounds what it can see:

  • An absence claimed in wording the lists do not contain is scored ignored, not misreport, so the published rate is a lower bound.
  • The lists cover English, Hindi and Hinglish. Another language falls out as not applicable rather than being guessed at.
  • The patterns are versioned (PATTERNS_VERSION) and recorded on every score, so two numbers from different lists are never compared as if they measured the same thing.
  • A judge fallback for unreadable answers is designed but not yet wired in, so today an answer the patterns cannot read leaves the denominator rather than being judged.

The lists are in src/neverempty/scorers/failure.py and are 30 lines of regex. If you only want the metric and not the harness, that file and the @tool wrapper are the parts worth copying.

The CI gate

The gate fails a build for three reasons only: a must_pass case failed, a paired test shows a significant regression against the committed baseline, or a hard floor was breached. Latency, cost and small accuracy wobbles are reported as warnings. That restraint is deliberate: a gate that fires on noise gets disabled within a month, which leaves you with no gate at all.

Exit Meaning
0 pass
1 significant regression (exact McNemar) or a floor breached
2 a must_pass case failed
3 inconclusive: too many cases unstable across repeats; rerun or reduce noise
4 invalid input: incomplete report, version mismatch, missing baseline

Codes 1, 2 and 3 fail the build. Code 3 says "rerun" rather than "regression", because a run too noisy to attribute a delta must not be read as the agent getting worse. Code 4 is an infrastructure failure and is labelled as such.

The paired test is one-sided: it compares outcomes per case on the same frozen set, so an improvement can never fail a build. At n=210, a real 5-point drop and noise are distinguishable; at n=100 they are not, which is why the sample size drives the design rather than the reverse.

Below ~200 cases, set floors or the gate detects almost nothing

Exact one-sided McNemar needs 5 clean pass-to-fail flips before it can reach p < 0.05 at all, because with zero improvements the tail is 0.5 ** flips: four flips give p = 0.0625 and pass. That floor is a property of the test, not of your suite, so it does not shrink as n shrinks — it just becomes a larger share of it:

Suite size 4 regressions pass silently which is
30 yes 13.3 points
60 yes 6.7 points
100 yes 4.0 points
210 yes 1.9 points
330 yes 1.2 points

This is the test behaving correctly — refusing to call four flips a regression is exactly the honesty the gate is for — but on a small suite it means the statistical gate alone will not catch a real drop. Set a hard floor, which is checked independently of the paired test:

[gate]
floors = { route = 0.95 }

floors defaults to empty, so until you set one, a suite under roughly 200 cases has a regression gate that fires only on large breaks. On 60 cases a 6.7-point drop passes at exit 0; with floors = { route = 0.95 } the same run exits 1 and names the metric. Use must_pass on individual cases for the requirements that must never break regardless of aggregate accuracy.

A noisy agent cannot pass, by design

Exit 3 is inconclusive, not regression: the run was too noisy to attribute any delta to the change. With the default 3 repeats and max_unstable_rate = 0.10, that fires at roughly 3.5% per-call nondeterminism:

Per-call noise Cases unstable Verdict
1% 3.1% passes
2% 5.8% passes
3.5% 10.2% exit 3
10% 26.9% exit 3

An agent above that line fails every build, including unchanged reruns, which is correct — a delta measured through that much noise means nothing — but it is not something to sit with. Fix the noise rather than the threshold:

  • temperature=0 and a fixed seed on every model call.
  • Stub the nondeterministic dependency. A tool whose ordering varies run to run will do this on its own.
  • Raise max_unstable_rate only deliberately, and treat every number from that suite as correspondingly noisier.

More repeats do not help: they measure the instability more precisely, they do not reduce it.

What the renderer refuses to print

Numbers in a README come from a committed report, because the renderer reads nothing else. It also refuses three things outright:

  • A metric with applicable=0 prints not measured, never 0%.
  • Below n=10 the percentage is suppressed and the raw counts shown instead.
  • Below n=50 the percentage is printed but labelled indicative.

A genuinely zero rate still prints as 0.0%, because 0% misreport is a result and the best possible one. It has to stay distinguishable from "we did not check".

The judge

Some checks string logic cannot make: whether an answer's reasoning contradicts a rule trace, or whether it implies absence without saying so. Those go to an LLM judge, which is the least trustworthy component here and is treated that way.

Bias Control
Self-preference The judge's model family must differ from the agent's. Enforced at construction and re-checked against the model the run resolves to. Fails closed on an unknown family.
Verbosity and framing The verifier sees one atomic claim, never the answer's tone, length, or the other claims.
Leniency on ambiguity Three labels, so the judge is never forced to pick supported or contradicted for something the evidence does not settle.
Instruction leakage Claim and evidence are delimited data with their closing tags escaped, and the prompt says instructions inside them are ignored. Four injection fixtures ship with the library.
Nondeterminism Temperature 0, a content-addressed cache so reruns are identical, and a measured self-consistency rate.
Invented labels Malformed output is retried twice and then labelled judge_error, never guessed.

A judge number is never published without its agreement figure. neverempty judge calibrate reports Cohen's kappa against human labels, the 3x3 matrix, and precision and recall for contradicted specifically. Below kappa 0.6 the judge-derived numbers are cut and only the deterministic checks are published.

Kappa rather than raw agreement, because raw agreement is inflated by the base rate: on a set that is 80% supported, a judge that always answers supported scores 80% agreement and kappa 0.

Core ships the JudgeModel protocol and a deterministic offline fake, never a provider SDK. A real binding is a dozen lines behind an extra, so a vendor outage never becomes a red build on unrelated work.

Numbers

Every number below is rendered from a committed report and links back to it. Each carries its sample size and 95% interval.

suite metric value 95% CI n notes
nextrole.routing calibration 89.7% 88.9% – 90.3% 330 bootstrap
nextrole.routing route 91.2% 87.7% – 93.8% 330 wilson

Reproducibility

  • nextrole.routing (test, suite version 1) — report, run 2026-10-02T08:53:16.454Z, target unrecorded, models unrecorded, pricing openai-2026-09-01, neverempty 0.1.0.

Every number will show its sample size, confidence interval, resolved model snapshot, date and git sha, and will link to the committed report it was rendered from. Judge-derived numbers are cut when kappa is below 0.6.

The region above is generated; regenerate and check it with:

neverempty readme evals/reports/*.json
neverempty readme evals/reports/*.json --check README.md   # CI

Documentation

Page Read it when
Getting started Going from an untraced tool to a CI gate.
Writing labels Before your first label. Decides whether your numbers mean anything.
Reading a report What each number is allowed to claim.
Architecture Why a piece is not simpler. Each answer is a failure mode.

Contributing

See CONTRIBUTING.md. Security policy in SECURITY.md.

License

Apache-2.0. See LICENSE.

Metadata

Release files for neverempty 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for neverempty 0.1.0
File Size Uploaded
neverempty-0.1.0.tar.gz 380.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for neverempty 0.1.0
File Interpreter ABI Platform
neverempty-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 575.9 kB

Release files / neverempty-0.1.0.tar.gz

Download URL neverempty-0.1.0.tar.gz
Size 380.3 kB
Tags Source
SHA-256 checksum
How to use checksums
c86827632e71f01cc6f6ed663975089753ca2eb979310a1dd11e020d485077ac
BLAKE2b-256 checksum
How to use checksums
a1e88c0607ed8fed12ee5836e14c9e98a5046d78c2689206c764a83ee725f5bb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / neverempty-0.1.0-py3-none-any.whl

Download URL neverempty-0.1.0-py3-none-any.whl
Size 195.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e1db833cdc6e52011ad9510633cc74f4bf0f5ac49ee78149f3c38676d749896f
BLAKE2b-256 checksum
How to use checksums
eff22a82854d1aadc8038a8b18bc2e3516c8e5ac326f27edff4c4de6da4e62cf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page