judgements
TypeSafe's System One model answers narrow, typed questions about a piece of state and returns
calibrated probabilities instead of text. judgements wraps that API so it looks like pydantic:
you declare the questions as fields of a model, and you get a model back.
hello.py is the whole idea in forty lines. This README explains what each line does.
Install
pip install judgements
To work on the library itself, run uv sync in this directory instead. It installs the package in
editable mode.
Run hello.py
Put your key in the environment:
export TYPESAFE_API_KEY=...
Then:
python hello.py
Triage(billing=True p=0.98, tone=frustrated p=1.00, urgency=today p=0.98)
route to billing
The model is the questions
class Triage(Judgements):
billing: bool = question("Is this ticket about billing?")
tone: Tone = question("What is the customer's tone in `body`?")
urgency: Urgency = score("How urgent is this ticket?")
triage = Triage.ask(ticket)
Each field is one question. The annotation is the answer type, the marker holds the instructions,
and Triage.ask(state) sends every field in a single request and returns a Triage whose fields
hold the plain answers.
There are three kinds of question, and the annotation picks the kind:
| Annotation | Kind | What comes back |
|---|---|---|
bool |
noul, "does this hold?" | True or False |
an Enum or Literal["a", "b"] |
choice, "which one?" | the chosen member |
an Enum with score(...) |
score, "how much, on this scale?" | the most probable level |
question(...) infers noul or choice from the annotation. Nothing in a type says "ordered rubric",
so a score is always declared with score(...). noul(...) and choice(...) exist when you want
to be explicit, and noul takes options: true= and false= describe the two outcomes, and
threshold= sets where the bool flips, 0.5 by default.
For an Enum, member names are the labels the model chooses between and member values are their descriptions. Write the descriptions as the concrete situations you mean:
class Urgency(Enum):
can_wait = "No deadline is implied; handle in the normal queue"
this_week = "The customer expects a resolution within a few days"
today = "The customer is blocked or demands immediate action"
For a score the order matters: first member is level 0, the lowest. For a Literal, the strings
are undescribed labels in a choice and the level descriptions themselves in a score.
Reading the answer
The fields are plain values, so triage.tone == Tone.angry and if triage.billing: just work.
The probabilities behind them are one attribute away:
triage.p.billing # 0.98, probability that the answer is yes
triage.p.tone # {Tone.calm: 0.0, Tone.frustrated: 1.0, Tone.angry: 0.0}
triage.confidence.tone # 1.0, the model's reported confidence in the chosen option
triage.expected.urgency # 1.97, see below
triage.results["urgency"] # ScoreResult(can_wait 0.01, this_week 0.01, today 0.98; expected 1.97)
triage.usage # Usage(requests=1, input_tokens=312, output_tokens=48)
p is short for probability: one number for a bool field, a distribution for the others.
expected exists for score fields only. Levels have positions, can_wait is 0, this_week is 1,
today is 2, and expected is the probability-weighted average of those positions. With the
probabilities above that is 0 × 0.01 + 1 × 0.01 + 2 × 0.98 = 1.97. It falls between levels and is
the number to use when averaging or ranking many items. The field itself holds the single most
probable level.
Use the probabilities to make policy explicit rather than trusting the top answer:
if triage.billing and triage.p.billing > 0.9:
route_to_billing()
elif triage.confidence.tone < 0.6:
escalate_to_human()
triage.model_dump() gives {'billing': True, 'tone': 'frustrated', 'urgency': 'today'}, with
Enum names rather than descriptions.
Because the answers and their probabilities share one object, a few field names are reserved:
p, confidence, expected, results, usage, questions, ask and from_answers. Using one
raises a TypeError at class definition.
Clients
Triage.ask(ticket) uses a default client that reads TYPESAFE_API_KEY. For anything beyond a
script, make a client:
ts = TypeSafe() # or AsyncTypeSafe(), then `await ts.ask(...)`
triage = ts.ask(ticket, Triage)
triage, refund = ts.ask(ticket, Triage, wants_refund) # several things, one request
triages = ts.map(tickets, Triage) # one request per ticket, in order
ts.usage # requests and tokens so far
ask accepts a pydantic model, a dict, a list or a string as state. It is sent as JSON exactly as
it is, so backticked paths in instructions, like `body`, are relative to the state itself.
Put related material together in one object when a judgement needs to compare parts of it.
ask also takes model=, retry= and timeout= for one call. The async map takes
concurrency=, eight by default.
Questions on their own
A question does not need a model. On its own it returns the full result object:
wants_refund = question("Does the customer explicitly ask for money back?", bool)
r = ts.ask(ticket, wants_refund) # NoulResult(no, p=0.08)
bool(r), r.probability
Options can be decided per request, which is how you rerank or select among candidates:
best = choice("Which of `candidates` best answers `query`?", candidates) # a list of labels
r = ts.ask({"query": query, "candidates": candidates}, best)
r.choice, r.ranked # the winner, and every candidate by probability
relevance = score("How relevant is `text` to `query`?", {"none": "Off topic", "partial": "Related", "direct": "Answers it"})
r = ts.ask({"query": query, "text": text}, relevance)
r.level, r.score, r.at_least("partial")
A dict gives each label a description. A list gives labels only.
Writing good questions
- Ask one narrow judgement per field. Split independent dimensions into separate fields; they are answered in parallel in the same request at no extra latency.
- Put the judgement in the instructions and the possible answers in the type. The field name is for your code and is not shown to the model.
- Include a way out when nothing may fit, such as an
unclearorothermember. - Check the exact request before spending tokens:
request(ticket, Triage) # {"state": {...}, "questions": {"Triage.billing": {...}, ...}}
Testing without a key
from judgements.testing import FakeTypeSafe
fake = FakeTypeSafe(billing=0.9, tone=Tone.angry, urgency={Urgency.today: 0.7, Urgency.this_week: 0.3})
triage = fake.ask(ticket, Triage) # same parsing path as the real client, no network
fake.requests[0]["state"] # what would have been sent
Answers are matched by field name. A bool field takes a probability or a bool; a choice or score
field takes the chosen option or a dict of option to probability. A missing answer raises. There is
an AsyncFakeTypeSafe too, and tests/test_judgements.py shows both in use:
python -m unittest discover -s tests
Further reading
The live docs are the reference for the model itself: System One, state, the three primitives and confidence.
Metadata
Release files for judgements 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| judgements-0.2.0.tar.gz | 16.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| judgements-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 34.0 kB
Release files / judgements-0.2.0.tar.gz
| Download URL | judgements-0.2.0.tar.gz |
|---|---|
| Size | 16.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
df7a542205065c11ef06cc589dc7ac971c35309856e9ed6ff39d0d6ff15994d1
|
|
BLAKE2b-256 checksum How to use checksums |
c1acac02c4de57dfbcda7d36f738c90ca670a7444d3afd580d3a47727952720d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|
Release files / judgements-0.2.0-py3-none-any.whl
| Download URL | judgements-0.2.0-py3-none-any.whl |
|---|---|
| Size | 17.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
aa7fbce67e02230adf1b91d63ee6b84e4494b49ea6eb652e1c9be0f91da09f7b
|
|
BLAKE2b-256 checksum How to use checksums |
211e1799620b7726926d0c68af94acafcd4b1d239d0f3dfad4b63129534fcf53
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|