onus
A claim checker. It finds the verifiable claims in a deck, memo or document and checks each one against the evidence. Numbers and dates are checked deterministically; everything else goes to a model that has to quote its source. Findings use ruff-style rule codes.
Status: alpha (0.1). Rule codes, config keys and the report format may change between releases. The deterministic tier is stable and tested; the model tiers depend on the model you point them at and will vary run to run.
$ onus update.pptx --evidence facts.md
update.pptx:slide 1: ONS002 Q1 2025: no date like it in the evidence
Q1 2025
update.pptx:slide 3: ONS001 95%: no number like it in the evidence
The stations holding renewal above 95% are the same ones that fixed bikes fastest.
update.pptx:slide 3: ONS002 18 months: no date like it in the evidence
Median repair backlog fell from 2.1x to 1.4x of weekly capacity over 18 months
update.pptx:slide 3: ONS002 2023: no date like it in the evidence
The shift followed the 2023 flood closures.
update.pptx:slide 4: ONS101 The evidence says nothing about why riders keep coming back; it only
provides operational and financial facts. p=0.97
Riders keep coming back because the stations sit beside daily commute routes.
8 findings across 17 claims (ONS002 ×6, ONS001 ×1, ONS101 ×1; numbers + model).
(Output trimmed to five of the eight findings.)
The onus is on the claim. It's the third of a set: riff lints the writing, overset lints the layout, and onus checks whether what the document says is true to its sources.
Why
An agent built that deck from a prompt that stated a handful of facts. Every figure above is one it invented: dates, periods, a retention threshold, a reason. The prose linter passed every one of them, because each read well. What they lacked was a source, and that's checkable.
How it works
onus is a small test pipeline. The claims are the test cases, the evidence is the fixture, and the checks are the assertions.
- Read the document into blocks: text frames per slide, paragraphs, table cells.
- Extract claims from two sources, merged:
- a deterministic scan finds every quantity: money (
$2.4B,2.4 billion,$2,400Mare all the same value), percentages, "two thirds", multiples, counts, years, quarters (Q1 2025), dates, and durations; - a pydantic-ai agent returns typed claims, adding the ones with no number in them ("customers are deepening usage because…") and what each claim is about. Its spans must appear in the document verbatim or they're dropped, its numbers are re-parsed from the text rather than trusted, and any number it missed becomes a claim of its own.
- a deterministic scan finds every quantity: money (
- Tier 1: deterministic. Each quantity is matched against every quantity
in the evidence, at the precision it was written with (
$2.4Bmatches$2,412M;42%doesn't match41%). It also matches the changes and ratios between quantities stated in one evidence sentence, so "1.4x, down from 2.1x" supports "down 33%". A model can't overrule this tier. - Tier 2: judged. A model sees a claim and the most relevant evidence passages, and answers supported, contradicted or unsupported. Its quote must appear in the evidence verbatim, or the verdict is discarded. It may judge claims with no number, and may clear a small count of listed things ("three theses" when three are listed). It can't clear a number, date or large count that tier 1 failed.
Evidence
onus deck.pptx --evidence facts.md --evidence data.json
onus memo.docx --evidence source_deck.pptx
onus deck.pptx --evidence ~/.harness/sessions/2026-10-01.jsonl # an agent transcript
onus deck.pptx --evidence-text "ARR $2.4B; NRR 118%"
onus deck.pptx --evidence ~/src/app --evidence-exclude 'tests/*' # a repository
Evidence can be .md/.txt, .json/.jsonl (flattened to text), .pptx/.docx,
or a Messages-API-shaped agent transcript. A transcript counts only what came
into the conversation: the user's messages and the results of tools that fetch
information. Excluded are the agent's own text, <system-reminder> blocks, and
tools that mostly echo the agent's own work back to it (read, write,
edit, bash, ls, find, grep, configurable via
transcript-exclude-tools). Otherwise a fabrication becomes its own evidence
the moment the agent reads back the file it wrote.
A directory
A directory, typically a git repository for a deck about the code in it, is
read as its files, each named by its path (app/README.md#3). Three things
keep a whole repository from drowning the claims in noise:
- What's read. In a git repository, only tracked files, so build output
and anything gitignored never count. Elsewhere, a walk that skips
.git,node_modules,.venv, caches and the like. Only known text types (prose, data, code),.pptx/.docx, andREADME-style names are read; lockfiles, minified bundles, binaries and files over 256 KB are skipped.--evidence-exclude GLOBskips more ('tests/*','*.csv'). A note says what was read:evidence: 138 tracked files from my-app/ (skipped 161 other types, 1 lockfiles). - What the judge sees. The six passages it gets per claim are ranked by BM25, so a rare word counts more than a common one, with code weighted at half of prose. A README paragraph beats the code that implements it.
- When a number counts. A repository contains every small number there
is, so a number read from a directory supports a claim only if its own
sentence (its line, in code and tables, plus a table's header) shares a
word with the claim.
MAX_TURNS = 40supports "the loop stops after 40 turns"; "113 unsupported claims" doesn't support "113 tool calls per deck". No arithmetic is derived from code. A file named on its own keeps the looser rule: any matching number counts.
Measured on a workshop deck checked against its own repository: of six made-up numbers, five matched something in the repo under the looser rule; under this one none did, and the one true number still passed. On the deck's 42 real numbered claims, one, a chart label "Build 6", went from "supported" (by "drifted on six axes") to flagged, which was correct.
Rules
| Code | Checked by | Rule |
|---|---|---|
| ONS001 | exact | A number the evidence does not contain or imply |
| ONS002 | exact | A year, quarter, date or duration the evidence does not contain |
| ONS003 | exact, model may clear | A small count (12 or under) the evidence does not state |
| ONS004 | exact | A number supported only by arithmetic (off by default; "show your working") |
| ONS101 | model | A factual claim the evidence does not support |
| ONS102 | model | A claim the evidence contradicts |
| ONS201 | model | A number the extractor missed (off by default; a measure of the extractor) |
Exit codes: 0 every claim supported, 1 findings, 2 usage or file error.
--no-llm runs tier 1 only: offline, deterministic, free. Without it onus
needs a provider key (ANTHROPIC_API_KEY for the default model) and stops
with a message, rather than silently degrading, if the key is missing.
--report claims.json writes the fact-check report: every claim, its
status, and the evidence passage that settled it. That's useful to a person
reviewing the deck, not only to the agent that wrote it.
Evals
evals/cases.json holds labelled cases from real agent runs. Each case lists
what must be flagged and what must not be.
$ uv run python evals/run.py # numbers only
case recall false + claims
deck: bike-share annual update 7/7 0 16
memo: planted fabrications 4/4 0 9
$ uv run python evals/run.py --model # the full pipeline
deck: bike-share annual update 9/9 0 19
memo: planted fabrications 4/4 0 9
The deterministic layer is held to perfect recall and no false positives on
these cases, as a test (tests/test_cli.py), because it's the part allowed
to refuse an answer. The model tier is measured, not guaranteed: it can miss a
claim with no number in it (an older Sonnet missed one on this deck), and
results vary by model and run.
Limits
- Magnitude, not direction. "Down 33%" and "up 50%" are both derivable from 1.4x and 2.1x; tier 1 doesn't know which way the metric moved.
- Presence, not subject. A number is supported if a named evidence file contains it anywhere (from a directory, its sentence must share a word with the claim). "NRR 42%" passes against "ARR up 42%". Catching a right number on the wrong subject is tier 2's job, and it currently judges only claims that tier 1 can't settle.
- Claims with a number and a qualitative part ("118% NRR tells us expansion is structural") are settled by their number.
Configuration
onus.toml (or .onus.toml, or [tool.onus] in pyproject.toml), found
from the document's folder upward:
select = ["ONS0", "ONS1"]
ignore = ["ONS003"]
extend-select = ["ONS004"]
llm = true
model = "anthropic:claude-sonnet-5-5" # any pydantic-ai model string
judge-threshold = 0.7
concurrency = 8
transcript-exclude-tools = ["read", "write", "edit", "bash", "ls", "find", "grep"]
Development
uv sync
uv run pytest -q --cov # hermetic: the model is pydantic-ai's TestModel/FunctionModel
uv run python evals/run.py # the labelled eval, offline
uv run ruff check src tests evals
uv run mypy # types
See CONTRIBUTING.md.
License
MIT.
Metadata
Release files for onus 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| onus-0.1.0.tar.gz | 154.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| onus-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 186.4 kB
Release files / onus-0.1.0.tar.gz
| Download URL | onus-0.1.0.tar.gz |
|---|---|
| Size | 154.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
91941c4468151ee4744b529ad51d94d87441857671c67cef89b02d4c7fd5ad4a
|
|
BLAKE2b-256 checksum How to use checksums |
4fc05d0873bb17b2a8381aeaa7a56e4ad9ddbd9715cd81f6a400453402a65e0c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / onus-0.1.0-py3-none-any.whl
| Download URL | onus-0.1.0-py3-none-any.whl |
|---|---|
| Size | 31.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
cb059a23808aefbd7ecc19ab34e6c44befcd50de389c559b9a354eb64f0d95a6
|
|
BLAKE2b-256 checksum How to use checksums |
700048a1765dcb840a53669bb1c7559d74b8f14468c75c44dd4f1903a6b00856
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|