quaesitor-zero
On TPC-DS — a public standard schema, with the data in front of it — a frontier model answered 3 of 10 questions that have no answer. Every one of those answers ran, was well formatted, and carried nothing to mark it as unsupported. It also declined 2 of 10 questions it could answer.
Does your data assistant say "I don't know" when it cannot know?
uvx quaesitor-zero generate --schema schema.sql --out questions.csv
# ask your assistant the questions, paste each response into the CSV
uvx quaesitor-zero score --answers questions.csv --out scorecard.html
Nothing is sent anywhere. The tool has no network code, no telemetry, and no model access — it reads your schema, writes twenty questions, and reads the answers back.
What it measures
A structurally unanswerable question is one the data cannot support: the attribute is not recorded, the period is outside the data, the two tables share no key. The correct response is to say so. An assistant that instead produces a confident, well-formatted, plausible number has failed silently — nothing downstream carries a signal that the number should not be trusted.
A tool that only measured refusals would be broken, because an assistant that refuses everything would score perfectly. So every run mixes two classes and scores the 2×2:
| Assistant answered | Assistant declined | |
|---|---|---|
| Unanswerable — correct action: decline | silent overreach | correct refusal |
| Answerable — correct action: answer | correct answer | over-refusal |
The headline is a pair, not a rate, and the summary statistic is balanced accuracy over the discrimination task, reported with coverage and risk at coverage in the vocabulary of selective prediction.
The matched answerable controls are not padding. They are what makes the unanswerable half mean anything.
The eight families
Every question comes from a family that is decidable from the schema, or from the schema plus a profile of the data. Nobody has to agree with us about what revenue means for the question to be unanswerable.
| # | Family | Derived from |
|---|---|---|
| 1 | Absent attribute | the column list |
| 2 | Out-of-range period | min/max of date columns |
| 3 | Missing grain | the foreign-key graph |
| 4 | Structurally absent value | a data profile |
| 5 | Unjoinable relation | components of the FK graph |
| 6 | Unmeasurable metric | absence of a required table |
| 7 | Ambiguous by construction | several columns for one business word |
| 8 | Absent population | distinct values of a filter column |
FAMILIES.md is the specification: what each family is, how the generator decides, and where each one can be wrong. It is the intellectual content of this tool, and it is worth reading before quoting a number from it.
Families 4 and 8 need a read-only connection; they emit nothing from DDL alone. Family 7 is off by default — it is the closest of the eight to a question about definitions, and the correct response to it is to ask, not to decline, so it is scored as a third outcome.
How it reaches your assistant
Through you. There is no adapter, no API key, no connection.
quaesitor-zero generate → questions.csv (id, question, response)
questions.key.json (which are unanswerable, and why)
You ask your assistant the questions however you normally would — the web UI is
fine — and paste each response into the response column. Then score reads
them back.
This works with an assistant that has only a UI, which is most of them. It needs no credentials and no security review, because nothing is connected. And the person running it reads every answer, which is the point: a silent failure that arrives as a summary statistic is a number, and one you read yourself is a problem.
The key file is separate on purpose. Putting expected: decline in the same
row as the question is one copy-paste away from your assistant's context, and an
assistant told which questions are traps scores well for a reason that has
nothing to do with the system being measured.
How responses are classified
Mechanically, by published rules, and never by a model. A tool whose central claim is that confident model output should not be trusted without a signal cannot rest its own headline number on a model's judgement.
Every classification prints the rule that fired. A response carrying evidence of
two different things — declining and also producing a figure — is not guessed
at: it is marked unclear, and score stops and asks you to read those
yourself before it will produce a scorecard. The count of them appears on the
page.
Rules are English by default and replaceable with --lexicon rules.json.
The scorecard
One self-contained HTML file: the 2×2 with Wilson intervals, balanced accuracy, coverage, risk at coverage, every question with its response and its classification, and a run fingerprint — schema digest, question-set digest, generator version, timestamp, counts, and whatever you called the assistant.
It also states its own boundary, because the boundary is the honest part:
This measures whether the system declines what it cannot answer. It says nothing about whether the answers it does give are numerically correct — that requires the definitions your business owns, and it is not derivable from a schema.
The worked example
make example # or: see examples/tpcds/README.md
Runs the whole thing against the TPC-DS standard schema, which is checked in, so it reproduces with no warehouse of your own and no model access.
What this is not
- not a correctness test — that needs your definitions, and it is the layer above
- not an eval platform — no accounts, no dashboards, no experiment tracking
- not a hosted service — having nothing to host is the feature
- not an adapter zoo — CSV first; adapters if somebody asks
- no telemetry of any kind, ever, including anonymous usage statistics
Install
pip install quaesitor-zero # also: uv tool install quaesitor-zero · pipx install quaesitor-zero
uvx quaesitor-zero --help # or run it without installing
uvx --from git+https://github.com/sandovabarbora/quaesitor-zero quaesitor-zero # latest from source
On PyPI: Python 3.10+, one dependency (DuckDB), Apache-2.0.
If a question is wrong
A generated question that is actually answerable makes a correct answer look like overreach, and it looks exactly like a real finding. Every question prints its warrant — the specific reason the schema cannot support it — so the mistake is findable. Telling us about one is the most useful thing anyone can do.
Part of Quaesitor — independent audit of AI answers over a data warehouse.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file quaesitor_zero-0.1.2.tar.gz.
File metadata
- Download URL: quaesitor_zero-0.1.2.tar.gz
- Upload date:
- Size: 871.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5813cdba09cbccf20fe96981420b6a230ed2e27bf2d883bb9abe8d130a00ba51
|
|
| MD5 |
5371292c5a7ae7d004b184e6719dd17f
|
|
| BLAKE2b-256 |
fc3a468b42de412d1202958d67f8255c00735d3ad602d58f8f0744a90a000afd
|
File details
Details for the file quaesitor_zero-0.1.2-py3-none-any.whl.
File metadata
- Download URL: quaesitor_zero-0.1.2-py3-none-any.whl
- Upload date:
- Size: 64.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.6 {"installer":{"name":"uv","version":"0.11.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f3b6904f7d5e232afdaae8017ac1dbc36aeacaafd2ea173fd7b31eaab8d1be2d
|
|
| MD5 |
8b3628566c49b09c2e1dcb49065c6365
|
|
| BLAKE2b-256 |
5782426ff211e7c9cd921f1f0ec44c1f3273b716b792ad0a69655341114aa7d1
|