Skip to main content

judgements

TypeSafe's System One model answers narrow, typed questions about a piece of state and returns calibrated probabilities instead of text. judgements wraps that API so it works with Pydantic models: you declare the questions as fields of a model, and you get a model back.

See hello.py for it in action.

Quickstart

pip install judgements

Run hello.py

You will need to set TYPESAFE_API_KEY in your environment.

python hello.py
Triage(billing=True p=0.98, tone=frustrated p=1.00, urgency=today p=0.98)
route to billing

The pydantic model is the set of questions

class Triage(Judgements):
    billing: bool = question("Is this ticket about billing?")
    tone: Tone = question("What is the customer's tone in `body`?")
    urgency: Urgency = score("How urgent is this ticket?")

triage = Triage.ask(ticket)

Each field poses a question, and the annotation sets the answer type. Triage.ask(state) sends every field in a single request and returns a Triage whose fields hold the answers.

There are three kinds of question, depending on the annotation.

Annotation Kind What comes back
bool noul, "does this hold?" True or False
an Enum or Literal["a", "b"] choice, "which one?" the chosen member
an Enum with score(...) score, "how much, on this scale?" the most probable level

question(...) infers noul or choice from the annotation.

For an Enum, member names are the labels and member values are their descriptions. Write descriptions as the concrete situations you mean:

class Urgency(Enum):
    can_wait = "No deadline is implied; handle in the normal queue"
    this_week = "The customer expects a resolution within a few days"
    today = "The customer is blocked or demands immediate action"

For a score, order matters: the first member is level 0. A Literal gives labels without descriptions.

Reading the answer

Fields are plain values, so triage.tone == Tone.angry and if triage.billing: just work. The probabilities are one attribute away:

triage.p.billing            # 0.98, probability that the answer is yes
triage.p.tone               # {Tone.calm: 0.0, Tone.frustrated: 1.0, Tone.angry: 0.0}
triage.confidence.tone      # 1.0, confidence in the chosen option
triage.expected.urgency     # 1.97, probability-weighted level from 0 to 2
triage.results["urgency"]   # ScoreResult(can_wait 0.01, this_week 0.01, today 0.98; expected 1.97)
triage.usage                # Usage(requests=1, input_tokens=312, output_tokens=48)

expected exists for score fields. It is the probability-weighted average of the level positions, and the number to use when ranking or averaging many items. The field itself holds the most probable level.

Use the probabilities to make policy explicit:

if triage.billing and triage.p.billing > 0.9:
    route_to_billing()
elif triage.confidence.tone < 0.6:
    escalate_to_human()

triage.model_dump() gives {'billing': True, 'tone': 'frustrated', 'urgency': 'today'}.

The names p, confidence, expected, results, usage, questions, ask and from_answers are reserved and cannot be fields.

Clients

Triage.ask(ticket) uses a default client that reads TYPESAFE_API_KEY. For anything beyond a script, make one:

ts = TypeSafe()                                   # or AsyncTypeSafe(), then `await ts.ask(...)`

triage = ts.ask(ticket, Triage)
triage, refund = ts.ask(ticket, Triage, wants_refund)      # several things, one request
triages = ts.map(tickets, Triage)                          # one request per ticket, in order
ts.usage                                                   # requests and tokens so far

State can be a pydantic model, a dict, a list or a string. It is sent as JSON as is, so backticked paths in instructions, like `body`, refer into the state.

ask also takes model=, retry= and timeout=. The async map takes concurrency=, eight by default.

Questions on their own

A question does not need a model. On its own it returns the full result:

wants_refund = question("Does the customer explicitly ask for money back?", bool)
r = ts.ask(ticket, wants_refund)      # NoulResult(no, p=0.08)
bool(r), r.probability

Options can be decided per request, which is how you rerank or select among candidates. A list gives labels, a dict adds descriptions:

best = choice("Which of `candidates` best answers `query`?", candidates)
r = ts.ask({"query": query, "candidates": candidates}, best)
r.choice, r.ranked                    # the winner, and every candidate by probability

relevance = score("How relevant is `text` to `query`?", {"none": "Off topic", "partial": "Related", "direct": "Answers it"})
r = ts.ask({"query": query, "text": text}, relevance)
r.level, r.score, r.at_least("partial")

Writing good questions

  • Ask one narrow judgement per field. Fields are answered in parallel in the same request.
  • Put the judgement in the instructions and the possible answers in the type. The field name is not shown to the model.
  • Include a way out when nothing may fit, such as an unclear or other member.
  • Check the exact request with request(ticket, Triage) before spending tokens.

Testing without a key

from judgements.testing import FakeTypeSafe

fake = FakeTypeSafe(billing=0.9, tone=Tone.angry, urgency={Urgency.today: 0.7, Urgency.this_week: 0.3})
triage = fake.ask(ticket, Triage)     # same parsing path as the real client, no network
fake.requests[0]["state"]             # what would have been sent

Answers are matched by field name: a probability or bool for a bool field, an option or a dict of option to probability for the others. AsyncFakeTypeSafe does the same for async code.

Further reading

The TypeSafe docs cover the model itself: System One, state, the three primitives and confidence.

Metadata

Release files for judgements 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for judgements 0.2.1
File Size Uploaded
judgements-0.2.1.tar.gz 15.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for judgements 0.2.1
File Interpreter ABI Platform
judgements-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 32.6 kB

Release files / judgements-0.2.1.tar.gz

Download URL judgements-0.2.1.tar.gz
Size 15.8 kB
Tags Source
SHA-256 checksum
How to use checksums
0107a7e534bdd07ed9ebd1949768ee8960c8172a54a43958cc62f8472841d6f4
BLAKE2b-256 checksum
How to use checksums
f45f0eee2622c195293d9749110b533b4123a60b5b21a043cb650d3cf27963de
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release files / judgements-0.2.1-py3-none-any.whl

Download URL judgements-0.2.1-py3-none-any.whl
Size 16.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4490fa152254092a08602111e44431241945cd4965ac57dd1d6b81f54901a3ba
BLAKE2b-256 checksum
How to use checksums
94817a1108bb96ff4ddedebfcef29144d2ffc0632b6e5e713942a33a3e090115
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page