neverempty
A Python eval harness that measures how often your tool-calling agent tells a user "no results" when a tool actually failed.
For engineers shipping agents who want that rate as a CI-gated number rather than an assumption. It injects tool faults on purpose, scores the answers, and fails the build when the misreport rate rises.
Status: 0.1.0, not released. Every milestone in the design doc is implemented, the suite is green on Python 3.10-3.13, and the Numbers section below carries a first real measurement: 91.2% routing accuracy over 330 hand-verified cases against a live gpt-4o router.
What is still missing is the number this library exists to produce.
misreport_as_emptyneeds injected faults, and those need the agent's tools wrapped with@tool; that work is in progress. Until it lands, the measurement machinery is demonstrated and the premise -- that agents report tool failures as absence often enough to matter -- is not.
Installation
Not on PyPI yet, so pip install neverempty does not work. Install from a
clone:
git clone https://github.com/Priyank032/neverempty
cd neverempty
pip install -e .
Core depends on pydantic>=2 and nothing else (plus tomli on Python 3.10,
where tomllib is not stdlib). Python 3.10 to 3.13. Integrations ship as
extras: neverempty[langgraph], neverempty[openai], neverempty[bedrock].
Start here
neverempty init # writes a working eval setup
neverempty run evals/neverempty.toml # runs it
That is a complete eval before you have written anything: three routing cases,
two with an injected tool failure, a target stub, and two reports. The numbers
in them describe the stub agent that init writes, not yours -- they show
the shape of the output, nothing more.
Then make it yours: point evals/target.py at your real agent, and replace the
example- cases with your own. neverempty coverage evals/neverempty.toml
says how many more each branch needs before a number from it is worth
publishing.
The cases are yours to write, like tests are. A label generated by a model
would make every number you publish a measure of one model agreeing with another, which is the opposite of what this is for.
The bug this exists to prevent
A tool times out. It returns []. The model reads [], concludes there is no
data, and tells the user "there are no results for Q3". The user believes it.
Nothing in the stack logged an error the user would ever see.
A tool wrapped in @tool cannot produce that shape. Three things change:
- The type. A tool returns
Ok,EmptyorErr— a tagged union discriminated onstatus, never a bare value whose emptiness you have to guess at. - The text the model reads. An error renders as an explicit instruction not
to claim absence, because a perfectly typed error that serialises to
[]in the tool message reproduces the bug exactly. - The measurement. Faults are injected on purpose, so the rate at which the agent misreports failure as absence is measured rather than assumed to be zero. Real timeouts are too rare in a test run to measure by waiting.
Quickstart
Install from a clone (PyPI release pending):
git clone https://github.com/Priyank032/neverempty
cd neverempty && pip install -e .
1. A tool that cannot lie about being empty
@tool is always called with parentheses. empty_when is the only way a
plain value becomes Empty — falsiness is never inferred, because 0, False
and "" are legitimate values.
from neverempty import tool
@tool(empty_when=lambda rows: rows == [])
async def search_jobs(city: str) -> list[dict]:
if city == "Nowhere":
return [] # -> Empty
return [{"title": "Backend Engineer", "city": city}] # -> Ok
@tool()
async def flaky_search(city: str) -> list[dict]:
raise TimeoutError("upstream timed out") # -> Err, never raises
The three outcomes stay distinguishable all the way into the prompt. This is
the whole point of the library, and it is what to_model() renders:
to_model() returns the JSON string to put in the tool message, not a
dict, so it goes straight into the message body with no further encoding:
print((await search_jobs(city="Pune")).to_model())
# {"status": "ok", "data": [{"title": "Backend Engineer", "city": "Pune"}], "truncated": false}
print((await search_jobs(city="Nowhere")).to_model())
# {"status": "empty", "note": "The query succeeded and returned no matching records."}
print((await flaky_search(city="Pune")).to_model())
# {"status": "error", "kind": "timeout", "note": "The tool failed. You do not know
# whether matching data exists. Do not say that no data exists."}
A failed call tells the model, in words, not to claim absence. A perfectly typed
error that serialises to [] in the tool message reproduces the original bug
exactly, so the type and the text are fixed together.
2. Your agent, as a plain async callable
The runner's target is async (Case, Tracer) -> None. No framework required —
the LangGraph adapter is a helper that builds such a callable, not a dependency.
async def run_agent(case, tracer):
query = case.input.messages[-1].content
result = await search_jobs(city="Pune" if "Pune" in query else "Nowhere")
if result.status == "error":
answer = "The job search failed, so I could not check."
elif result.status == "empty":
answer = "No jobs found."
else:
answer = f"Found {len(result.value)} job(s)."
tracer.current_run.set_output(answer=answer, route="job_search")
3. A dataset
One JSON object per line. Every field in expect is optional, and a scorer
whose expectation is absent reports not applicable — never a pass.
{"schema_version":1,"id":"demo-0001","suite":"demo.routing","split":"dev",
"input":{"messages":[{"role":"user","content":"jobs in Pune"}]},
"expect":{"route":{"label":"job_search"}}}
4. Run, report, gate
import asyncio
from neverempty import Dataset, Runner, Tracer, gate, render_markdown, scorers
from neverempty.report.gate import GateConfig
from neverempty.tracer.sinks import JsonlSink
dataset = Dataset.load("demo.jsonl") # fails on any bad line, with line numbers
runner = Runner(
target=run_agent,
scorers=[scorers.route()],
tracer=Tracer(sink=JsonlSink("traces.jsonl")),
repeats=3, # instability is reported, not hidden
concurrency=4,
seed=20260929, # the harness's randomness; see the note below
)
report = asyncio.run(runner.run(dataset))
report.save("report.json")
print(render_markdown(report))
result = gate(report, report, config=GateConfig())
raise SystemExit(result.exit_code)
That run produces:
dataset: 3 cases validated
report: status=ok complete=True cases=3 scored=3
metric route: value=1.0 n=3 ci=(0.439, 1.0)
traces: 9 written (3 cases x 3 repeats)
gate: verdict=pass exit=0
seed fixes the harness's own randomness: sampling, the bootstrap, and the
order work is scheduled in. It cannot reach the agent's. The same seed with a
nondeterministic agent still gives different scores run to run — set
temperature=0 and a fixed seed on your model calls for that, and see the
instability ceiling below.
The interval is wide because n=3. That is the point: the renderer refuses to print a percentage below n=10, and labels anything below n=50 as indicative.
5. The CLI
neverempty run evals/neverempty.toml # run every suite the config declares
neverempty validate evals/**/*.jsonl # schema, duplicate ids, split hash
neverempty coverage evals/neverempty.toml # per-branch label backlog; non-zero if short
neverempty compare base.json cand.json # paired stats, markdown diff
neverempty gate base.json cand.json # exit 1 on a real regression
neverempty render report.json # markdown report
neverempty import traces.jsonl --cases cases.jsonl # traces from another language
neverempty judge calibrate labels.jsonl # kappa, 3x3 matrix, per-language slice
neverempty readme evals/reports/*.json # published numbers, linked to their reports
Exit codes from gate are the contract: 0 pass, 1 regression, 2 must-pass
failed, 3 inconclusive, 4 invalid.
If you only want the idea
Most of this repository is an eval harness: a runner, a dataset schema, a tracer, a judge, a statistics package, a CI gate. If you already have an eval stack you like — Opik, DeepEval, promptfoo, Langfuse, Inspect — you do not need any of that, and bolting a second tracer and dataset format onto a working setup is a bad trade.
The part that is actually novel is small. Four files, about 740 lines including docstrings:
| File | What it is |
|---|---|
core/results.py |
the Ok / Empty / Err union and what each renders to a model |
core/classify.py |
mapping an exception to an error kind, with cancellation left alone |
core/faults.py |
injecting a tool failure without touching the tool |
scorers/failure.py |
the 30 regexes that decide misreport vs reported vs ignored |
Copy those into your own stack, write one scorer that reads your traces, and you have the metric. That is a genuinely reasonable thing to do, and it is faster than adopting this.
What you get by using the harness instead is the parts that are easy to get
subtly wrong: any-hit collapse across repeats, a metric that reads "not
measured" instead of 0 when nothing was measured, a gate that refuses an
incomplete run rather than treating it as a smaller sample, and intervals that
are not printed below n=10. Those are the bugs this repository's own history is
mostly made of — the measurement layer is harder to get right than the idea.
Protocol-level error flags already exist and are worth using regardless: MCP's
isError, Anthropic's is_error, LangChain's ToolMessage.status="error".
None of them distinguishes empty from failed, which is the distinction this
is about, but all of them beat swallowing an exception into [].
Design principles
Four rules hold everywhere in the codebase:
- Failure is never representable as empty. No falsiness inference:
0,Falseand""are legitimate values. - Missing never looks like zero. An unknown cost is
nullwith a reason, never0. - Not-applicable never looks like failure. A scorer that cannot decide
returns
None, neverpassed=False. - Cancellation propagates untouched.
What 0.1.0 contains
Typed tool results and the @tool wrapper; a contextvars tracer writing a
versioned JSON trace format, which an agent in another language can emit to be
scored (there is no SDK for one); a dataset schema; an async runner with
repeats, budget caps and record/replay; deterministic scorers plus a judge and
a judge calibrate command reporting Cohen's kappa against your labels; pure
Python statistics (Wilson intervals, exact McNemar, seeded bootstrap); and a
CI gate that fails on a one-sided exact McNemar test at p<0.05, a breached
floor, or a failed must-pass case, and exits 3 when more than 10% of cases are
unstable across repeats.
Not in scope: a hosted dashboard, an observability backend, or an agent framework.
Also deferred past 0.1.0: neverempty.langchain.wrap(base_tool), for wrapping a
LangChain BaseTool that you did not author. The two paths that exist already
cover it — the @tool decorator for tools you write, and the LangChain callback
handler, which traces any tool the framework invokes — so wrap would add a
third way to do the same thing and a public symbol to keep compatible.
How repeats collapse
Two rules, deliberately asymmetric, because one rule would be wrong in both directions at once:
- Capability metrics (route, first-tool, all-args, argument accuracy) collapse by majority over repeats. A single flake is not a broken capability. A tie resolves to failure.
- Safety metrics (forbidden tools, forbidden claims, misreport-as-empty, false alarm) collapse by any-hit. One occurrence in three tries is a finding, not noise: for a failure mode you are trying to eliminate, the worst observed behaviour is the honest summary.
- Fact recall is continuous, so it takes the median of the values.
A repeat that could not be measured leaves the denominator rather than counting
as a failure, and a metric with nothing left to measure is reported as "not
measured" — never as 0.
How misreport_as_empty is scored
This is the number the library exists to produce, so here is exactly how it is computed and where it is weak.
After a run injects a tool fault, the agent's final answer is matched against
two regex lists: 13 absence patterns (there are no, could not find any,
koi ... nahi) and 17 failure patterns (search failed, could not check, तकनीकी समस्या). The answer is then labelled:
| Label | Condition | Counts as |
|---|---|---|
misreport |
an absence pattern fired | a hit — the headline number |
reported |
a failure pattern fired, and no absence pattern | a pass |
ignored |
neither fired | a fail, but not a misreport |
| not applicable | neither fired and the answer is in a script the lists do not cover | excluded, with a reason |
An absence claim beats a failure mention: an answer that says both "the search
failed" and "there are no jobs" is a misreport, because the user acts on the
absence.
misreport_as_empty is the share of fault cases labelled misreport, with a
Wilson interval. Across repeats it is any-hit: one lie in three runs makes the
case a hit, because a lie that happens a third of the time is still a lie.
What this means for the number. It is a deterministic pattern match, not a model's judgment. That makes it reproducible and free, and it bounds what it can see:
- An absence claimed in wording the lists do not contain is scored
ignored, notmisreport, so the published rate is a lower bound. - The lists cover English, Hindi and Hinglish. Another language falls out as not applicable rather than being guessed at.
- The patterns are versioned (
PATTERNS_VERSION) and recorded on every score, so two numbers from different lists are never compared as if they measured the same thing. - A judge fallback for unreadable answers is designed but not yet wired in, so today an answer the patterns cannot read leaves the denominator rather than being judged.
The lists are in src/neverempty/scorers/failure.py and are 30 lines of regex.
If you only want the metric and not the harness, that file and the @tool
wrapper are the parts worth copying.
The CI gate
The gate fails a build for three reasons only: a must_pass case failed, a
paired test shows a significant regression against the committed baseline, or a
hard floor was breached. Latency, cost and small accuracy wobbles are reported as
warnings. That restraint is deliberate: a gate that fires on noise gets disabled
within a month, which leaves you with no gate at all.
| Exit | Meaning |
|---|---|
| 0 | pass |
| 1 | significant regression (exact McNemar) or a floor breached |
| 2 | a must_pass case failed |
| 3 | inconclusive: too many cases unstable across repeats; rerun or reduce noise |
| 4 | invalid input: incomplete report, version mismatch, missing baseline |
Codes 1, 2 and 3 fail the build. Code 3 says "rerun" rather than "regression", because a run too noisy to attribute a delta must not be read as the agent getting worse. Code 4 is an infrastructure failure and is labelled as such.
The paired test is one-sided: it compares outcomes per case on the same frozen set, so an improvement can never fail a build. At n=210, a real 5-point drop and noise are distinguishable; at n=100 they are not, which is why the sample size drives the design rather than the reverse.
Below ~200 cases, set floors or the gate detects almost nothing
Exact one-sided McNemar needs 5 clean pass-to-fail flips before it can
reach p < 0.05 at all, because with zero improvements the tail is
0.5 ** flips: four flips give p = 0.0625 and pass. That floor is a property
of the test, not of your suite, so it does not shrink as n shrinks — it just
becomes a larger share of it:
| Suite size | 4 regressions pass silently | which is |
|---|---|---|
| 30 | yes | 13.3 points |
| 60 | yes | 6.7 points |
| 100 | yes | 4.0 points |
| 210 | yes | 1.9 points |
| 330 | yes | 1.2 points |
This is the test behaving correctly — refusing to call four flips a regression is exactly the honesty the gate is for — but on a small suite it means the statistical gate alone will not catch a real drop. Set a hard floor, which is checked independently of the paired test:
[gate]
floors = { route = 0.95 }
floors defaults to empty, so until you set one, a suite under roughly 200
cases has a regression gate that fires only on large breaks. On 60 cases a
6.7-point drop passes at exit 0; with floors = { route = 0.95 } the same run
exits 1 and names the metric. Use must_pass on individual cases for the
requirements that must never break regardless of aggregate accuracy.
A noisy agent cannot pass, by design
Exit 3 is inconclusive, not regression: the run was too noisy to attribute
any delta to the change. With the default 3 repeats and
max_unstable_rate = 0.10, that fires at roughly 3.5% per-call
nondeterminism:
| Per-call noise | Cases unstable | Verdict |
|---|---|---|
| 1% | 3.1% | passes |
| 2% | 5.8% | passes |
| 3.5% | 10.2% | exit 3 |
| 10% | 26.9% | exit 3 |
An agent above that line fails every build, including unchanged reruns, which is correct — a delta measured through that much noise means nothing — but it is not something to sit with. Fix the noise rather than the threshold:
temperature=0and a fixed seed on every model call.- Stub the nondeterministic dependency. A tool whose ordering varies run to run will do this on its own.
- Raise
max_unstable_rateonly deliberately, and treat every number from that suite as correspondingly noisier.
More repeats do not help: they measure the instability more precisely, they do not reduce it.
What the renderer refuses to print
Numbers in a README come from a committed report, because the renderer reads nothing else. It also refuses three things outright:
- A metric with
applicable=0prints not measured, never0%. - Below n=10 the percentage is suppressed and the raw counts shown instead.
- Below n=50 the percentage is printed but labelled indicative.
A genuinely zero rate still prints as 0.0%, because 0% misreport is a result
and the best possible one. It has to stay distinguishable from "we did not check".
The judge
Some checks string logic cannot make: whether an answer's reasoning contradicts a rule trace, or whether it implies absence without saying so. Those go to an LLM judge, which is the least trustworthy component here and is treated that way.
| Bias | Control |
|---|---|
| Self-preference | The judge's model family must differ from the agent's. Enforced at construction and re-checked against the model the run resolves to. Fails closed on an unknown family. |
| Verbosity and framing | The verifier sees one atomic claim, never the answer's tone, length, or the other claims. |
| Leniency on ambiguity | Three labels, so the judge is never forced to pick supported or contradicted for something the evidence does not settle. |
| Instruction leakage | Claim and evidence are delimited data with their closing tags escaped, and the prompt says instructions inside them are ignored. Four injection fixtures ship with the library. |
| Nondeterminism | Temperature 0, a content-addressed cache so reruns are identical, and a measured self-consistency rate. |
| Invented labels | Malformed output is retried twice and then labelled judge_error, never guessed. |
A judge number is never published without its agreement figure. neverempty judge calibrate reports Cohen's kappa against human labels, the 3x3 matrix, and
precision and recall for contradicted specifically. Below kappa 0.6 the
judge-derived numbers are cut and only the deterministic checks are published.
Kappa rather than raw agreement, because raw agreement is inflated by the base
rate: on a set that is 80% supported, a judge that always answers supported
scores 80% agreement and kappa 0.
Core ships the JudgeModel protocol and a deterministic offline fake, never a
provider SDK. A real binding is a dozen lines behind an extra, so a vendor outage
never becomes a red build on unrelated work.
Numbers
Every number below is rendered from a committed report and links back to it. Each carries its sample size and 95% interval.
| suite | metric | value | 95% CI | n | notes |
|---|---|---|---|---|---|
| nextrole.routing | calibration |
89.7% | 88.9% – 90.3% | 330 | bootstrap |
| nextrole.routing | route |
91.2% | 87.7% – 93.8% | 330 | wilson |
Reproducibility
- nextrole.routing (test, suite version 1) — report, run 2026-10-02T08:53:16.454Z, target
unrecorded, models unrecorded, pricingopenai-2026-09-01, neverempty0.1.0.
Every number will show its sample size, confidence interval, resolved model snapshot, date and git sha, and will link to the committed report it was rendered from. Judge-derived numbers are cut when kappa is below 0.6.
The region above is generated; regenerate and check it with:
neverempty readme evals/reports/*.json
neverempty readme evals/reports/*.json --check README.md # CI
Documentation
| Page | Read it when |
|---|---|
| Getting started | Going from an untraced tool to a CI gate. |
| Writing labels | Before your first label. Decides whether your numbers mean anything. |
| Reading a report | What each number is allowed to claim. |
| Architecture | Why a piece is not simpler. Each answer is a failure mode. |
Contributing
See CONTRIBUTING.md. Security policy in SECURITY.md.
License
Apache-2.0. See LICENSE.
Metadata
Release files for neverempty 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| neverempty-0.1.0.tar.gz | 380.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| neverempty-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 575.9 kB
Release files / neverempty-0.1.0.tar.gz
| Download URL | neverempty-0.1.0.tar.gz |
|---|---|
| Size | 380.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c86827632e71f01cc6f6ed663975089753ca2eb979310a1dd11e020d485077ac
|
|
BLAKE2b-256 checksum How to use checksums |
a1e88c0607ed8fed12ee5836e14c9e98a5046d78c2689206c764a83ee725f5bb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / neverempty-0.1.0-py3-none-any.whl
| Download URL | neverempty-0.1.0-py3-none-any.whl |
|---|---|
| Size | 195.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e1db833cdc6e52011ad9510633cc74f4bf0f5ac49ee78149f3c38676d749896f
|
|
BLAKE2b-256 checksum How to use checksums |
eff22a82854d1aadc8038a8b18bc2e3516c8e5ac326f27edff4c4de6da4e62cf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log