Skip to main content

judgements

TypeSafe's System One model answers narrow, typed questions about a piece of state and returns calibrated probabilities instead of text. judgements wraps that API so it looks like pydantic: you declare the questions as fields of a model, and you get a model back.

hello.py is the whole idea in forty lines. This README explains what each line does.

Install

From this directory:

uv sync

This installs judgements in editable mode. Plain pip install -e . works too if you prefer your own environment.

Run hello.py

Put your key in the environment:

export TYPESAFE_API_KEY=...

Then:

python hello.py
Triage(billing=True p=0.98, tone=frustrated p=1.00, urgency=today p=0.98)
route to billing

The model is the questions

class Triage(Judgements):
    billing: bool = question("Is this ticket about billing?")
    tone: Tone = question("What is the customer's tone in `body`?")
    urgency: Urgency = score("How urgent is this ticket?")

triage = Triage.ask(ticket)

Each field is one question. The annotation is the answer type, the marker holds the instructions, and Triage.ask(state) sends every field in a single request and returns a Triage whose fields hold the plain answers.

There are three kinds of question, and the annotation picks the kind:

Annotation Kind What comes back
bool noul, "does this hold?" True or False
an Enum or Literal["a", "b"] choice, "which one?" the chosen member
an Enum with score(...) score, "how much, on this scale?" the most probable level

question(...) infers noul or choice from the annotation. Nothing in a type says "ordered rubric", so a score is always declared with score(...). noul(...) and choice(...) exist when you want to be explicit, and noul takes options: true= and false= describe the two outcomes, and threshold= sets where the bool flips, 0.5 by default.

For an Enum, member names are the labels the model chooses between and member values are their descriptions. Write the descriptions as the concrete situations you mean:

class Urgency(Enum):
    can_wait = "No deadline is implied; handle in the normal queue"
    this_week = "The customer expects a resolution within a few days"
    today = "The customer is blocked or demands immediate action"

For a score the order matters: first member is level 0, the lowest. For a Literal, the strings are undescribed labels in a choice and the level descriptions themselves in a score.

Reading the answer

The fields are plain values, so triage.tone == Tone.angry and if triage.billing: just work. The probabilities behind them are one attribute away:

triage.p.billing            # 0.98, probability that the answer is yes
triage.p.tone               # {Tone.calm: 0.0, Tone.frustrated: 1.0, Tone.angry: 0.0}
triage.confidence.tone      # 1.0, the model's reported confidence in the chosen option
triage.expected.urgency     # 1.97, see below
triage.results["urgency"]   # ScoreResult(can_wait 0.01, this_week 0.01, today 0.98; expected 1.97)
triage.usage                # Usage(requests=1, input_tokens=312, output_tokens=48)

p is short for probability: one number for a bool field, a distribution for the others.

expected exists for score fields only. Levels have positions, can_wait is 0, this_week is 1, today is 2, and expected is the probability-weighted average of those positions. With the probabilities above that is 0 × 0.01 + 1 × 0.01 + 2 × 0.98 = 1.97. It falls between levels and is the number to use when averaging or ranking many items. The field itself holds the single most probable level.

Use the probabilities to make policy explicit rather than trusting the top answer:

if triage.billing and triage.p.billing > 0.9:
    route_to_billing()
elif triage.confidence.tone < 0.6:
    escalate_to_human()

triage.model_dump() gives {'billing': True, 'tone': 'frustrated', 'urgency': 'today'}, with Enum names rather than descriptions.

Because the answers and their probabilities share one object, a few field names are reserved: p, confidence, expected, results, usage, questions, ask and from_answers. Using one raises a TypeError at class definition.

Clients

Triage.ask(ticket) uses a default client that reads TYPESAFE_API_KEY. For anything beyond a script, make a client:

ts = TypeSafe()                                   # or AsyncTypeSafe(), then `await ts.ask(...)`

triage = ts.ask(ticket, Triage)
triage, refund = ts.ask(ticket, Triage, wants_refund)      # several things, one request
triages = ts.map(tickets, Triage)                          # one request per ticket, in order
ts.usage                                                   # requests and tokens so far

ask accepts a pydantic model, a dict, a list or a string as state. It is sent as JSON exactly as it is, so backticked paths in instructions, like `body`, are relative to the state itself. Put related material together in one object when a judgement needs to compare parts of it.

ask also takes model=, retry= and timeout= for one call. The async map takes concurrency=, eight by default.

Questions on their own

A question does not need a model. On its own it returns the full result object:

wants_refund = question("Does the customer explicitly ask for money back?", bool)
r = ts.ask(ticket, wants_refund)      # NoulResult(no, p=0.08)
bool(r), r.probability

Options can be decided per request, which is how you rerank or select among candidates:

best = choice("Which of `candidates` best answers `query`?", candidates)   # a list of labels
r = ts.ask({"query": query, "candidates": candidates}, best)
r.choice, r.ranked                    # the winner, and every candidate by probability

relevance = score("How relevant is `text` to `query`?", {"none": "Off topic", "partial": "Related", "direct": "Answers it"})
r = ts.ask({"query": query, "text": text}, relevance)
r.level, r.score, r.at_least("partial")

A dict gives each label a description. A list gives labels only.

Writing good questions

  • Ask one narrow judgement per field. Split independent dimensions into separate fields; they are answered in parallel in the same request at no extra latency.
  • Put the judgement in the instructions and the possible answers in the type. The field name is for your code and is not shown to the model.
  • Include a way out when nothing may fit, such as an unclear or other member.
  • Check the exact request before spending tokens:
request(ticket, Triage)               # {"state": {...}, "questions": {"Triage.billing": {...}, ...}}

Testing without a key

from judgements.testing import FakeTypeSafe

fake = FakeTypeSafe(billing=0.9, tone=Tone.angry, urgency={Urgency.today: 0.7, Urgency.this_week: 0.3})
triage = fake.ask(ticket, Triage)     # same parsing path as the real client, no network
fake.requests[0]["state"]             # what would have been sent

Answers are matched by field name. A bool field takes a probability or a bool; a choice or score field takes the chosen option or a dict of option to probability. A missing answer raises. There is an AsyncFakeTypeSafe too, and tests/test_judgements.py shows both in use:

python -m unittest discover -s tests

Further reading

The live docs are the reference for the model itself: System One, state, the three primitives and confidence.

Metadata

Release files for judgements 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for judgements 0.1.0
File Size Uploaded
judgements-0.1.0.tar.gz 16.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for judgements 0.1.0
File Interpreter ABI Platform
judgements-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 33.9 kB

Release files / judgements-0.1.0.tar.gz

Download URL judgements-0.1.0.tar.gz
Size 16.4 kB
Tags Source
SHA-256 checksum
How to use checksums
4e8d8c668c62f4df6811291ed6d73bc0e50d6634d26a0a0b84608a8432f0e5b8
BLAKE2b-256 checksum
How to use checksums
079c98a5ce674adef3d594b9fc9c56beeb5655062860fc16a6fd586c414b2ef6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release files / judgements-0.1.0-py3-none-any.whl

Download URL judgements-0.1.0-py3-none-any.whl
Size 17.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7c28c599f8f5956230d791872b2b2cdae1fb86e13018d3e7a5ec7530df169410
BLAKE2b-256 checksum
How to use checksums
f1d48df522f38fe8977ab486d9bf63aa175750f675f8eb9abfaa9530457a7568
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release history Release notifications | RSS feed

0.2.1

2 release files

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page