Skip to main content

Semantic Operators

pip install "semantic-operators[typesafe]"

Write a semantic judgment once, run it on any System One model, and get a "don't know" instead of a guess when the model isn't sure.

System One models are small, fast models that classify instead of generating text, as opposed to a chat LLM. You ask them typed questions (yes/no, pick one, rate on a scale) about text or data, and they return answers with probabilities. Jev was the first; Laya is an open-weight, Jev-compatible alternative you can run locally. More are coming.

Provider Models Where it runs Install
providers.typesafe.TypeSafe Jev (jev-latest, or pin a version) TypeSafe's hosted API (needs TYPESAFE_API_KEY) [typesafe]
providers.laya.Laya Laya checkpoints (English, multilingual, typed-decisions) on your machine (~800 MB download on first use) [laya]

A provider is the service or runtime you talk to; the model is a setting.

Install

pip install "semantic-operators[typesafe]"        # TypeSafe (hosted Jev)
pip install "semantic-operators[laya]"            # Laya (local; pulls in torch)
pip install "semantic-operators[typesafe,laya]"   # both

The core alone (pip install semantic-operators) has no dependencies.

The whole idea

A System One model is asked named questions about a piece of state and returns an answer with probabilities for each. There are three kinds of question:

Question You give it answer.value
Boolean instructions (+ optional true/false meanings) True / False
Choice instructions + named options the chosen option name
Score instructions + ordered rubric levels expected level as a float, e.g. 1.7

Every Answer also carries probabilities (a dict, in the question's option/level order), confidence (the probability of its own answer), raw (the provider's own answer object), and call (which model answered, below).

A provider is anything with one method:

def ask(self, state, questions: dict[str, Question]) -> dict[str, Answer]

That's the entire abstraction.

Quick start

echo "TYPESAFE_API_KEY=..." > .env
uv run --env-file .env --extra typesafe python examples/hello.py
from typesafe_sdk import TypeSafeClient
from semantic_operators import Boolean, Choice, Score
from semantic_operators.providers.typesafe import TypeSafe

with TypeSafeClient() as client:          # you create and own the SDK client
    provider = TypeSafe(client)           # model defaults to "jev-latest"
    answers = provider.ask(
        "I was charged twice and I'm furious.",
        {
            "is_complaint": Boolean("Is the customer complaining?"),
            "department": Choice("Which team should handle this?",
                                 {"billing": "Payments, refunds", "other": "Anything else"}),
            "urgency": Score("How urgent is this?", ["low", "medium", "high"]),
        },
    )

answers["department"].value           # "billing"
answers["department"].probabilities   # {"billing": 0.97, "other": 0.03}

Swapping to Laya changes only how the provider is built:

import laya
from semantic_operators.providers.laya import Laya

provider = Laya(laya.load("convaiinnovations/laya"))   # or Laya(laya.Router())
answers = provider.ask(state, questions)                # same questions, same Answer type

Compare both side by side:

uv run --env-file .env --extra typesafe --extra laya python examples/compare.py

Named operators

This is where the library gets its name. An operator is a semantic judgment defined once, with a name, and used anywhere:

from semantic_operators import Boolean, Choice, Score
from semantic_operators.operators import Operator, apply

is_complaint = Operator("is_complaint", Boolean("Is the customer complaining?"))
urgency = Operator("urgency", Score("How urgent is this?", ["low", "medium", "high"]))

is_complaint(provider, message).value                        # one operator, one call
answers = apply(provider, message, [is_complaint, urgency])  # several, still one call
answers["urgency"].value
  • Combine freely. System One models answer many questions in one pass, so apply asks any set of operators in a single provider call. Names must be unique.
  • Provider-neutral. An operator doesn't hold a provider; you pass one in, so the same operator runs on TypeSafe, Laya, or anything else.
  • Wording is part of the operator. It changes the answers (see the benchmark), so keep operators in code, under version control, and benchmark them as they are. operators.questions([...]) turns them into the dict bench.run takes.
  • Async: await op.call_async(provider, state) and await apply_async(...).

examples/triage.py builds ticket triage from three operators, sending anything the model isn't sure about to a person. Add --laya to run the same code locally.

Reranking

Your retrieval (search, vector index, database) finds candidates; a System One model reorders them by how well each one answers the query. Each candidate is scored on its own against a relevance rubric, one call each, and plain code sorts the results:

from semantic_operators import Score
from semantic_operators.rerank import rerank, rerank_async, reweighted

relevance = Score("How useful is the document for answering the query?",
                  ["no useful information", "on topic but doesn't answer", "partly answers",
                   "answers with minor gaps", "fully answers"])

ranking = rerank(provider, relevance, query="How do I reset my password?",
                 candidates={"doc-1": {"title": ..., "text": ...}, ...})  # retrieval order
for r in ranking.top(10):
    r.id, r.score, r.answer.probabilities
  • Each call sees only {"query", "document"} (plus "context" if you pass one). Ids and positions are never sent, so a score doesn't depend on the other candidates.
  • Ranked by expected rubric level, not confidence. Ties keep retrieval order, and scores aren't normalized across documents: all can be relevant, or none.
  • Nothing gets a made-up score. Undecided or failed candidates go to unscored. top() refuses a partial ranking unless you pass allow_partial=True, and refuses scores from more than one model (ranking.models), since a moving alias such as jev-latest can change mid-run.
  • Your own level weights: reweighted(ranking, [0, 10, 40, 80, 100]) re-sorts from the probabilities already returned, with no new calls. Uneven weights can change the order, so evaluate them first.
  • Async: await rerank_async(..., concurrency=8) keeps up to 8 calls in flight.

bench.ndcg(order, grades) scores an ordering against graded relevance labels. examples/rerank.py reranks three hand-graded searches and compares NDCG before and after; on those, TypeSafe went from 0.56 (retrieval order) to 1.00 and Laya to 0.84.

"Don't know" answers

A model that's split, or not sure enough, should say so rather than guess. Give any question a min_confidence; below it, the answer comes back undecided (value is None, decided is False), with its probabilities kept:

department = Choice("Which team?", {"billing": None, "technical": None},
                    min_confidence=0.8)
answer = provider.ask(message, {"department": department})["department"]

if answer.decided:
    route(answer.value)
else:
    send_to_a_human(answer.probabilities)   # still shows what it was leaning toward

An exact tie (a Boolean at 0.5, two options equally likely) is always undecided. Confidence is the model's own view, not a guarantee. benchmarks/run_confidence.py checks whether it means anything: on the support suite, TypeSafe's urgency answers at min_confidence=0.8 were right 10 of 10 times (answering half the cases), while Laya's were right 4 of 7.

Which model answered

jev-latest moves over time, so every answer records what its provider call reported:

answers["department"].call
# Call(provider='TypeSafe', model='jev-1.13.0', input_tokens=291, output_tokens=20)

All answers from one call share one Call. model is exactly what the provider reported, which can differ from what you asked for (above, jev-latest). Laya reports a fixed agent name (laya-rl-agent), not which checkpoint answered. Anything a provider doesn't report is None.

Errors

Every provider raises one error type, whatever went wrong underneath (network, bad key, rate limit, a model failure, an unexpected response):

from semantic_operators import ProviderError

try:
    answers = provider.ask(state, questions)
except ProviderError as error:
    error.provider     # "TypeSafe" or "Laya"
    error.__cause__    # the original exception, for details

Provider output is checked before it becomes an Answer: probabilities must be finite, between 0 and 1, sum to 1 (allowing for the providers' rounding), and agree with the answer. A malformed response raises ProviderError rather than looking like a confident answer.

Retries and timeouts belong to the client you build. The TypeSafe SDK retries by default; to have every failure reach you (for example, when something above you does its own retrying), turn that off:

TypeSafeClient(retry=RetryPolicy(max_retries=0, timeout=10.0))
``` Questions check themselves too: a `Choice` needs at least two distinct options,
a `Score` at least two distinct levels, and a bad definition raises `ValueError`.

## Benchmark

`bench.run(provider, questions, cases)` asks each labeled case all questions in one call
and reports, per question, **accuracy** (Score values are rounded to the nearest level)
and **p(correct)**, the average probability the provider gave the right answer, plus
latency and every miss. Undecided answers are counted separately (`answered`), not as
misses. `bench.at_min_confidence(report, questions, cases, 0.8)` re-scores a run at a
stricter threshold without asking the provider again.

```sh
uv run --env-file .env --extra typesafe --extra laya python benchmarks/run.py

benchmarks/support_tickets.py holds 20 hand-written, hand-labeled support messages and the same 3 questions in three wordings. bench.stability(reports) reports how often a provider's decision stays the same when only the wording changes (labels play no part). It's a smoke test, not a verdict: small, authored, one person's labels.

Async

Every provider has an async twin with the same contract, await provider.ask(...):

from typesafe_sdk import AsyncTypeSafeClient
from semantic_operators.providers.typesafe import AsyncTypeSafe
from semantic_operators.providers.laya import AsyncLaya

async with AsyncTypeSafeClient() as client:
    answers = await AsyncTypeSafe(client).ask(state, questions)

AsyncLaya runs the local model in a worker thread, one call at a time. Concurrency speeds up a hosted API (many requests in flight), not a single local model. bench.run_async(provider, questions, cases, concurrency=8) benchmarks async providers:

uv run --env-file .env --extra typesafe --extra laya python benchmarks/run_async.py

Layout

src/semantic_operators/
  types.py          Boolean, Choice, Score, Answer, Call, make_answer: our vocabulary
  provider.py       Provider and AsyncProvider (one method each)
  errors.py         ProviderError, the one error every provider raises
  providers/typesafe.py  translates to/from the TypeSafe SDK
  providers/laya.py translates to/from the laya package
  operators.py      (higher layer) named operators: define once, combine in one call
  rerank.py         (higher layer) rerank search results by relevance
  bench.py          (higher layer) run labeled cases through a provider, score them;
                    ndcg for rankings
examples/
  hello.py          one real call to Jev
  compare.py        the same questions through TypeSafe and Laya
  triage.py         ticket triage built from named operators
  rerank.py         rerank three searches, NDCG before and after
benchmarks/
  support_tickets.py  20 labeled messages + the questions
  run.py              runs the suite through TypeSafe and Laya
  run_async.py        concurrency, and both providers at once
  run_confidence.py   the "don't know" trade-off at several min_confidence levels
tests/                offline tests (uv run --extra typesafe pytest); CI runs them
ROADMAP.md            where this is headed

Layers

Semantic Operators is built in layers inside one package:

  1. Base layer: a clean, provider-neutral abstraction over System One models: types.py, provider.py, errors.py, providers/.
  2. Higher layers: built only on the base layer: operators.py (named operators), rerank.py (reranking), and bench.py (benchmarking).

The base layer never imports from a higher layer, so it could later be split out as its own package without changing how it's used.

Rules

  • The library never reads API keys or environment variables. You build the client.
  • The core has no dependencies. Each provider's SDK is an optional extra ([typesafe], [laya]).
  • Our names, not the provider's: Boolean, not noul.

Not here yet (on purpose)

Operators are just named questions for now. Combining them is where this is headed: conditions ("ask B only when A says yes"), chains, and small decision flows built from operators. Also planned: a provider-neutral call timeout, and suites as data files. See ROADMAP.md.

License

MIT

Release files for semantic-operators 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for semantic-operators 0.6.0
File Size Uploaded
semantic_operators-0.6.0.tar.gz 102.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for semantic-operators 0.6.0
File Interpreter ABI Platform
semantic_operators-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 125.4 kB

Release files / semantic_operators-0.6.0.tar.gz

Download URL semantic_operators-0.6.0.tar.gz
Size 102.0 kB
Tags Source
SHA-256 checksum
How to use checksums
9684b66fcfa082e8d74579ecf38f5a153965c7766c21a8ac2e596ae78064e311
BLAKE2b-256 checksum
How to use checksums
afeede608c127aa0811f8fdbd3fd4347cd19f94fd6ecc4b55cfc9c9b75d1ff97
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / semantic_operators-0.6.0-py3-none-any.whl

Download URL semantic_operators-0.6.0-py3-none-any.whl
Size 23.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8421f0fa1d4e46d10f31f5fb7e869d15ebe0d38e1679acb73b37e4ef5ba9a518
BLAKE2b-256 checksum
How to use checksums
58a0c6c99837b96e26e14ab74de481f212162ca42718cbe0865d97d09b74e709
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.10.1

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page