Skip to main content

judgements

TypeSafe's System One model answers narrow, typed questions about a piece of state and returns calibrated probabilities instead of text. judgements wraps that API so it looks like pydantic: you declare the questions as fields of a model, and you get a model back.

hello.py is the whole idea in forty lines. This README explains what each line does.

Install

pip install judgements

To work on the library itself, run uv sync in this directory instead. It installs the package in editable mode.

Run hello.py

Put your key in the environment:

export TYPESAFE_API_KEY=...

Then:

python hello.py
Triage(billing=True p=0.98, tone=frustrated p=1.00, urgency=today p=0.98)
route to billing

The model is the questions

class Triage(Judgements):
    billing: bool = question("Is this ticket about billing?")
    tone: Tone = question("What is the customer's tone in `body`?")
    urgency: Urgency = score("How urgent is this ticket?")

triage = Triage.ask(ticket)

Each field is one question. The annotation is the answer type, the marker holds the instructions, and Triage.ask(state) sends every field in a single request and returns a Triage whose fields hold the plain answers.

There are three kinds of question, and the annotation picks the kind:

Annotation Kind What comes back
bool noul, "does this hold?" True or False
an Enum or Literal["a", "b"] choice, "which one?" the chosen member
an Enum with score(...) score, "how much, on this scale?" the most probable level

question(...) infers noul or choice from the annotation. Nothing in a type says "ordered rubric", so a score is always declared with score(...). noul(...) and choice(...) exist when you want to be explicit, and noul takes options: true= and false= describe the two outcomes, and threshold= sets where the bool flips, 0.5 by default.

For an Enum, member names are the labels the model chooses between and member values are their descriptions. Write the descriptions as the concrete situations you mean:

class Urgency(Enum):
    can_wait = "No deadline is implied; handle in the normal queue"
    this_week = "The customer expects a resolution within a few days"
    today = "The customer is blocked or demands immediate action"

For a score the order matters: first member is level 0, the lowest. For a Literal, the strings are undescribed labels in a choice and the level descriptions themselves in a score.

Reading the answer

The fields are plain values, so triage.tone == Tone.angry and if triage.billing: just work. The probabilities behind them are one attribute away:

triage.p.billing            # 0.98, probability that the answer is yes
triage.p.tone               # {Tone.calm: 0.0, Tone.frustrated: 1.0, Tone.angry: 0.0}
triage.confidence.tone      # 1.0, the model's reported confidence in the chosen option
triage.expected.urgency     # 1.97, see below
triage.results["urgency"]   # ScoreResult(can_wait 0.01, this_week 0.01, today 0.98; expected 1.97)
triage.usage                # Usage(requests=1, input_tokens=312, output_tokens=48)

p is short for probability: one number for a bool field, a distribution for the others.

expected exists for score fields only. Levels have positions, can_wait is 0, this_week is 1, today is 2, and expected is the probability-weighted average of those positions. With the probabilities above that is 0 × 0.01 + 1 × 0.01 + 2 × 0.98 = 1.97. It falls between levels and is the number to use when averaging or ranking many items. The field itself holds the single most probable level.

Use the probabilities to make policy explicit rather than trusting the top answer:

if triage.billing and triage.p.billing > 0.9:
    route_to_billing()
elif triage.confidence.tone < 0.6:
    escalate_to_human()

triage.model_dump() gives {'billing': True, 'tone': 'frustrated', 'urgency': 'today'}, with Enum names rather than descriptions.

Because the answers and their probabilities share one object, a few field names are reserved: p, confidence, expected, results, usage, questions, ask and from_answers. Using one raises a TypeError at class definition.

Clients

Triage.ask(ticket) uses a default client that reads TYPESAFE_API_KEY. For anything beyond a script, make a client:

ts = TypeSafe()                                   # or AsyncTypeSafe(), then `await ts.ask(...)`

triage = ts.ask(ticket, Triage)
triage, refund = ts.ask(ticket, Triage, wants_refund)      # several things, one request
triages = ts.map(tickets, Triage)                          # one request per ticket, in order
ts.usage                                                   # requests and tokens so far

ask accepts a pydantic model, a dict, a list or a string as state. It is sent as JSON exactly as it is, so backticked paths in instructions, like `body`, are relative to the state itself. Put related material together in one object when a judgement needs to compare parts of it.

ask also takes model=, retry= and timeout= for one call. The async map takes concurrency=, eight by default.

Questions on their own

A question does not need a model. On its own it returns the full result object:

wants_refund = question("Does the customer explicitly ask for money back?", bool)
r = ts.ask(ticket, wants_refund)      # NoulResult(no, p=0.08)
bool(r), r.probability

Options can be decided per request, which is how you rerank or select among candidates:

best = choice("Which of `candidates` best answers `query`?", candidates)   # a list of labels
r = ts.ask({"query": query, "candidates": candidates}, best)
r.choice, r.ranked                    # the winner, and every candidate by probability

relevance = score("How relevant is `text` to `query`?", {"none": "Off topic", "partial": "Related", "direct": "Answers it"})
r = ts.ask({"query": query, "text": text}, relevance)
r.level, r.score, r.at_least("partial")

A dict gives each label a description. A list gives labels only.

Writing good questions

  • Ask one narrow judgement per field. Split independent dimensions into separate fields; they are answered in parallel in the same request at no extra latency.
  • Put the judgement in the instructions and the possible answers in the type. The field name is for your code and is not shown to the model.
  • Include a way out when nothing may fit, such as an unclear or other member.
  • Check the exact request before spending tokens:
request(ticket, Triage)               # {"state": {...}, "questions": {"Triage.billing": {...}, ...}}

Testing without a key

from judgements.testing import FakeTypeSafe

fake = FakeTypeSafe(billing=0.9, tone=Tone.angry, urgency={Urgency.today: 0.7, Urgency.this_week: 0.3})
triage = fake.ask(ticket, Triage)     # same parsing path as the real client, no network
fake.requests[0]["state"]             # what would have been sent

Answers are matched by field name. A bool field takes a probability or a bool; a choice or score field takes the chosen option or a dict of option to probability. A missing answer raises. There is an AsyncFakeTypeSafe too, and tests/test_judgements.py shows both in use:

python -m unittest discover -s tests

Further reading

The live docs are the reference for the model itself: System One, state, the three primitives and confidence.

Metadata

Release files for judgements 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for judgements 0.2.0
File Size Uploaded
judgements-0.2.0.tar.gz 16.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for judgements 0.2.0
File Interpreter ABI Platform
judgements-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 34.0 kB

Release files / judgements-0.2.0.tar.gz

Download URL judgements-0.2.0.tar.gz
Size 16.5 kB
Tags Source
SHA-256 checksum
How to use checksums
df7a542205065c11ef06cc589dc7ac971c35309856e9ed6ff39d0d6ff15994d1
BLAKE2b-256 checksum
How to use checksums
c1acac02c4de57dfbcda7d36f738c90ca670a7444d3afd580d3a47727952720d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release files / judgements-0.2.0-py3-none-any.whl

Download URL judgements-0.2.0-py3-none-any.whl
Size 17.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
aa7fbce67e02230adf1b91d63ee6b84e4494b49ea6eb652e1c9be0f91da09f7b
BLAKE2b-256 checksum
How to use checksums
211e1799620b7726926d0c68af94acafcd4b1d239d0f3dfad4b63129534fcf53
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page