Skip to main content

solvi

PyPI CI License: Apache-2.0 Hugging Face

Build decision systems from a catalog of Python functions and checks plus typed questions, and get answers you can verify.

Why

You describe a task with plain Python functions (computations, checks, answer rules) and questions with typed answers (yes/no, a choice, a score, "not stated", a span of the text, a ranking, a number range). For each request, a strategist plans which functions and checks to run for the asked questions. Every answer comes with:

  • a confidence (calibratable per question);
  • a reason you can check: the rule inputs, a formula over computed facts, or a quote with character offsets in the source document;
  • a hash-chained trace of every step, which can be re-executed later to confirm the answer or pinpoint the step that was altered.

When something cannot be computed, a function fails, or a rule returns an answer outside the allowed options, solvi abstains instead of guessing. A failed hard check always overrides any model confidence.

Grounded decisions. Fuzzy proposes, deterministic decides, everything is in the trace. Each fact and answer records its provenance — given, computed, quoted, decided (a model's choice among options, with probabilities) or learned — and model-backed steps record the model's id and fingerprint, so a replay can tell when the model changed since a decision. A model's quote that is not literally the text at its offsets, or a choice outside its options, is rejected and counted (system.stats); a fallback producer runs or the question abstains. print(res.audit()) shows what each answer rests on, which safeguards fired, and how much of its support is deterministic. A decision without models and one with models are the same system (examples/12_grounded_audit.py).

And it is fast. The strategist plans a flow over a 10 000-part catalog in about 6 ms and runs only the parts the questions need (2.4% of that catalog). Hard checks run first, so a failing one skips the expensive rest; independent slow parts (API calls, model inference) run in parallel. On an insurance-claim desk with slow services (examples/09_strategy_at_scale.py) a full decision takes 463 ms instead of 1 122 ms for a script that computes everything, and 152 ms when an expired policy settles the claim first.

Try it

  • solvi playground: write a decision task in Python and run it, watch the strategist's plan, tamper with a trace and see the replay catch it, learn rules from examples.
  • solvi documents: cited, typed answers from contracts, invoices, receipts, leases and more — the ModernBERT extractor (ONNX) and the decisions both run in your browser; add a field by describing it.
  • solvi arcade: game agents that explain every move — tic-tac-toe, maze, minesweeper, 20 questions, Mafia detective, a bot arena, and "hack the trace".
  • solvi realms: an endless strategy game whose factions are solvi systems — tested for 100 000 turns: flat decision time (~0.3–0.6 ms), bounded memory and state, every sampled trace replay OK.
  • All run entirely in your browser (Pyodide): no server, no GPU, nothing you type leaves the page.
  • Models: solvi-ai/solvi-large (typed decisions, 396M, preview), solvi-ai/solvi-base (the same answers on a CPU / in ONNX, 150M, preview), solvi-ai/extract-base (fields by description) and solvi-ai/extract-receipts. Each model card states what the model was measured on, how well it does, and its limits; all models are listed at huggingface.co/solvi-ai.

gallery/ — twelve decision tasks across directions (support triage, email routing, content guard, security alerts, AI-agent audit, release rollout, KYC/AML, clinical screening, credit with adverse-action reasons, procurement 3-way match, double-charge refunds, predictive maintenance), each with scenarios, a runner and a side-by-side against an answer-only model.

Install

pip install solvi              # core: rules, checks, learned answer heads (numpy, scipy, pydantic)
pip install "solvi[model]"     # + torch, transformers: ModernBERT field extractors for documents and the decider
pip install "solvi[onnx]"      # + onnxruntime, tokenizers: the decider (solvi.decide) on CPU without torch

Requires Python 3.10+.

Quickstart (core only, no model)

from datetime import date
from solvi import Answer, Catalog, Question, System

cat = Catalog()

@cat.fn                                            # argument names = facts it reads; function name = fact it sets
def days_requested(start, end):
    return (end - start).days + 1

@cat.fn
def remaining_after(balance, days_requested):
    return balance - days_requested

@cat.check(hard=True, then={"approve": "reject"})  # if this check is False, "approve" is forced to "reject"
def enough_balance(remaining_after):
    return remaining_after >= 0

@cat.check
def enough_notice(start, today, days_requested):
    return days_requested < 5 or (start - today).days >= 14

@cat.rule("approve")
def approve(enough_notice):
    return "approve" if enough_notice else "needs_manager"

system = System(cat, [Question("approve", "Approve the leave?",
                               Answer.choice(["approve", "needs_manager", "reject"]),
                               checkpoints=["enough_balance"])])
res = system.ask({"start": date(2026, 10, 19), "end": date(2026, 10, 23),
                  "today": date(2026, 9, 25), "balance": 14})
print(res["approve"].answer, res["approve"].confidence, res["approve"].why)
print(res.computed_state)
print(res.trace.replay(cat))

Output:

approve 1.0 enough_notice = True
days_requested           = 5
remaining_after          = 9
enough_balance           = True
enough_notice            = True
{'ok': True, 'steps': 5, 'mismatches': [], 'models': []}

With "balance": 3 the hard check fails and the answer is reject with status == "forced", whatever the rule says. solvi.show.show(res, cat) prints answers, the planned flow, the computed state and the replay result in one go.

How it works

  • Catalog. @cat.fn (computation), @cat.check (bool), @cat.extract (value from text, returned as a Quote with offsets) and @cat.rule(question) (answer rule). A part's contract is its signature: argument names are the facts it reads, the function name is the fact it sets. Type hints are optional and become the facts' types (def risk_score(risk_points: dict[str, float]) -> float): producer and consumer types are checked when a part is registered, values are validated / coerced with pydantic at run time, and a value that fails is rejected like an ungrounded quote (safeguard type_rejected). Untyped parts cost nothing.
  • Questions. Question(name, text, Answer.yes_no() | Answer.choice([...]), checkpoints=[...]). Questions without a rule get a small answer head trained from labeled examples (system.fit) or a readable learned rule list (system.learn_rule).
  • Strategist. For each question it walks backwards from the rule's arguments (or the learned features) through the catalog signatures to the keys of init_state, adds the question's checkpoints and every check that touches a computed fact. Everything else in the catalog is not executed; the flow records why each part was taken or skipped.
  • Execution. Each part runs once, even if several questions need it. Hard checks and their inputs run first; when one fails, the steps only the settled questions needed are skipped (res.trace.skipped). With System(..., workers=8) independent steps run in parallel threads as soon as their inputs are ready. Results go into computed_state with their provenance; extracted values keep their quote. Each step record is hashed and chained to the previous one in flow order, so the trace does not depend on scheduling.
  • Answers and trace. res[q].answer / .confidence / .why / .status (ok, forced, abstain), plus res.trace.replay(catalog), which recomputes every step from recorded inputs and reports mismatches, broken hash links and quotes outside the text.

Typed decisions with a model

Types declare questions; the model proposes; checks decide. The fields of a pydantic model are the questions, their types the kinds (one option, several, an ordered score, yes/no); a decider (solvi.decide, solvi[onnx] or solvi[model]) answers them about a text or a JSON / pydantic state — several in one forward pass when the checkpoint can — with probabilities, a calibrated confidence and act / escalate. Hard checks, constraints and rules still decide.

from typing import Literal
from pydantic import BaseModel, Field
from solvi import Catalog, Scale, System
from solvi.decide import DecideModel

class Triage(BaseModel):
    team: Literal["billing", "technical", "shipping"] = Field(description="Which team should handle this ticket?")
    urgency: Scale[Literal["low", "medium", "high", "critical"]] = Field(description="How urgent is it?")
    angry: bool = Field(description="Is the customer angry?")
    topics: list[Literal["refund", "delay", "bug"]] = Field(description="What does the ticket mention?")

model = DecideModel.load("solvi-ai/solvi-base")         # or a local folder; solvi_decide.json says what it can do
cat = Catalog()
questions = model.questions(cat, Triage, text_fact="ticket", escalate_below=0.6)

@cat.check(hard=True, then={"urgency": "critical"})         # a legal threat is critical, whatever the model says
def no_legal_threat(ticket) -> bool:
    return "lawyer" not in str(ticket).lower()
questions[1].checkpoints.append("no_legal_threat")

res = System(cat, questions).ask({"ticket": {"subject": "Charged twice", "body": "Refund my double payment!",
                                             "customer": {"tier": "pro"}}})
print({q: (r.answer, r.status) for q, r in res.results.items()})
print(res.audit("team"))           # probabilities, the model's fingerprint, the shared pass, what escalated and why

A state is read as key paths (customer.tier: pro), the format the decider is trained on; an unsure or escalated answer abstains with the reason (system.stats["model_escalated"], ["low_confidence"]). The checkpoint contract is in docs/decide_format.md; examples/15_typed_decisions.py runs the whole story with a stand-in model. Published deciders are previews: read the model card before relying on one, and fit it on 30–60 labelled examples of your task (part.fit, part.calibrate_for) — checks, constraints and escalation are what make the answers safe to act on, not the model alone.

Every answer is a value and a confidence, and the types also declare answer primitives: Maybe[T] ("not stated" — solvi.Unknown — is a real answer, unlike an abstention), Span[float] (an exact piece of the text, parsed), Rank[...] (the top k, in order), Estimate[0, 7, 14] (a number with an interval), and evidence quotes on any answer (Claim(value, evidence=[...]), Question(require_evidence=True)) — each checked in the text, from rules or a model (guide, examples/16_primitives.py).

Planning around dead ends and costs (code strategist)

The default strategist needs the inputs of every producer of a fact. solvi.strategy.ModelStrategist() plans around producers whose inputs are never given, and with producers="equivalent" picks the cheapest verified plan by declared cost= (an exact 0/1 program; hard checks that govern a question always stay in the plan). No model is involved; the plan is one hashed record in the trace and replay re-verifies it.

from solvi.strategy import ModelStrategist
system = System(cat, questions, strategist=ModelStrategist(producers="equivalent"))

A segment model that proposes producers when costs are not declared, and solvi.aliases (wiring parameter names that match no fact), ship as experimental; their weights are not published. See docs/strategist.md and examples/17_model_strategist.py.

Extract from documents

With solvi[model], fields are found by a fine-tuned ModernBERT extractor. The extracted value is always a span of the document, so every answer built on it can be cited.

from solvi.extract_multi import MultiSpanExtractor

ex = MultiSpanExtractor(["total", "date"])          # ModernBERT-large, one pass per document for all fields
ex.fit(train_docs, train_spans, epochs=4)           # train_spans: [{"total": (start, end), "date": (start, end) | None}]
cat.extract(ex.field("total"))                      # doc -> Quote(text, start, end, confidence)
cat.extract(ex.field("date"))

@cat.fn
def amount(total):                                  # extracted values are strings; parse them in ordinary functions
    return float(total.replace(",", ""))

For long documents (contracts), solvi.extract_long.LongSpanExtractor reads the whole text in overlapping 1024-token windows, takes a field description instead of a fixed field list, and supports "no answer" with a per-field threshold. See docs/guide.md.

Results

Same training documents for both sides. The baseline, Laya, is a ModernBERT-large model that answers the typed questions directly, fine-tuned with its authors' recipe. Details and caveats: docs/benchmarks.md.

Task (test set) solvi Baseline
SROIE receipts, 6 questions (361 receipts) 97.8% 92.6%
SROIE, only 100 labeled training receipts 95.9% 80.0%
CORD receipts, 4 questions (100 receipts) 98.5% 96.0%
CORD, share of questions answered at >= 99% precision 99.7% 14.2%
CUAD contracts, 5 questions (102 contracts, median 33k chars) 94.8% 77.5% (sees first 1024 tokens only)
  • Calibrated confidence: ECE 0.008 (SROIE), 0.011 (CORD), 0.027 (CUAD).
  • On the document benchmarks, 100% of answers are backed by a quote at stated offsets, or the system abstains.
  • Speed: about 0.4 ms per decision when no model is involved; about 39 ms per receipt with the one-pass extractor on an A100 GPU.

Speed

Strategist on random layered catalogs (benchmarks/strategist_scale.py, one CPU core):

catalog parts plan plan + run + trace parts run share of catalog
100 0.2 ms 1.3 ms 31 31%
1 000 0.7 ms 2.6 ms 115 12%
10 000 6 ms 10 ms 236 2.4%

Insurance claim desk with six slow services of 100-300 ms (examples/09_strategy_at_scale.py):

questions asked script computing everything solvi, one by one solvi, workers=8
fast track? 1 122 ms 503 ms 313 ms
full decision + payout 1 122 ms 773 ms 461 ms
all five questions 1 122 ms 1 025 ms 463 ms
expired policy (hard check settles it) 1 122 ms 152 ms 153 ms

A rule-only decision on a small catalog takes well under a millisecond; on documents the extractor dominates (about 39 ms per receipt with the one-pass extractor on an A100).

When to use it

  • The answer is a computation or a rule over a few values found in a document or a record: receipts, invoices, contracts, requests, orders.
  • You need to show why: auditors, compliance, or a human reviewing low-confidence cases.
  • Some rules are non-negotiable (hard checks), and the rest can be learned from about 100 labeled examples.

When not to use it

  • Open-ended free-text questions or generated answers. solvi answers typed questions only: yes/no, choices, scores, multi-label, "not stated", exact spans of the text, rankings and number ranges.
  • New fields with no labeled examples. Extracting a field from its description alone does not work yet (14% and 66% on two held-out fields); a universal extractor is coming.
  • No labels at all. Plan on roughly 100 labeled documents (field positions) per task.
  • CPU-only deployment with a quantized model: the int8 ONNX extractor loses up to 12 points on amounts, company names and addresses. fp32 on CPU keeps accuracy but takes about 0.7 s per receipt on 2 cores.

Examples

File What it shows
examples/01_leave_request.py Leave request from a plain dict: rules, hard checks, parts the strategist skips
examples/02_shop_order.py Shop order: two rule-based questions plus a "suspicious?" question learned from history with fit
examples/03_invoices.py Invoice approval from text: regex extractors with quotes, four questions, one learned
examples/04_refunds.py Refund e-mails: yes/no learned from examples under a hard "within 30 days" check
examples/05_tic_tac_toe.py Tic-tac-toe agent: each move is an answer with its reason (win, block, fork, ...); a hard check rejects invalid boards
examples/06_learned_rules.py learn_rule: route parcels to delivery zones from free-form addresses with a readable if-then list learned from labels
examples/09_strategy_at_scale.py Insurance claim desk: the strategist generates a different plan per question set, hard checks first with early exit, slow services in parallel; timed
examples/10_learn_in_milliseconds.py fit_fast: a new question learned in milliseconds, then corrected one example at a time (each correction ~0.2 ms, nothing retrained)
examples/11_answer_types_and_constraints.py Multi-label and ordinal answers tied by constraints between answers; contradictions in learned answers are repaired by joint decoding
examples/12_grounded_audit.py One catalog with and without models: provenance, res.audit(), a hallucinated quote caught by grounding, a decision outside its options, a model changed since the decision, lifetime safeguard stats
examples/13_decide_model.py Support-email routing by a decider model as a catalog part: bias correction on unlabelled emails, 16 labelled examples, abstention, a constraint with a rule-based question, a hard check, the audit, teach, escalation for a target error rate, a JSON ticket (the real model with SOLVI_DECIDE_MODEL, a stand-in otherwise)
examples/14_typed_catalog.py Typed facts: a pydantic request, type hints as fact types, answer types from the rules' return types (Enum, Literal, bool), a mismatch caught at registration, rejected values → fallback / abstention, a response as JSON that loads back and replays
examples/15_typed_decisions.py Typed decisions: a pydantic ticket, the questions as a pydantic model's fields (choice, ordinal score, yes/no, multi-label), four answers from one forward pass, a hard check, a constraint and a rule over the model, an escalation in the audit and stats (the real model with SOLVI_DECIDE_MODEL, a stand-in otherwise)
examples/16_primitives.py Answer primitives: "not stated" vs abstain, spans parsed into numbers, evidence quotes checked in the text (require_evidence), a ranking with scores, an estimate with an interval — from rules and from a decider with the L14g contract; confidence per kind, JSON round trip, replay
examples/17_model_strategist.py The code strategist: dead ends dropped, the cheapest verified plan by declared costs, a model's proposal checked and rejected; aliases for names that match no fact (experimental; stand-ins without weights)
examples/07_receipts_model.py Expense check on a scanned receipt: a receipts-tuned extractor cites each field, rules and a hard check decide (needs solvi[model])
examples/08_contracts_by_description.py Contract review with fields defined only in words: the general extractor reads the whole contract, cites clauses or says "absent" (needs solvi[model])

Run them from a clone: python examples/01_leave_request.py.

More

License

Apache-2.0. See LICENSE.

Citation

A paper is in preparation. Until then, please cite the repository:

@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}

Metadata

Release files for solvi 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for solvi 0.5.1
File Size Uploaded
solvi-0.5.1.tar.gz 196.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for solvi 0.5.1
File Interpreter ABI Platform
solvi-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 404.5 kB

Release files / solvi-0.5.1.tar.gz

Download URL solvi-0.5.1.tar.gz
Size 196.9 kB
Tags Source
SHA-256 checksum
How to use checksums
b4e3a2f1964db197c6337cb5a07569c7388d6c46bcf8feb227588c5c2b028448
BLAKE2b-256 checksum
How to use checksums
c4ac3e51ea2b8045d4470aa053274521c78f8b3499efa02d7b40a4bc2defd42a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / solvi-0.5.1-py3-none-any.whl

Download URL solvi-0.5.1-py3-none-any.whl
Size 207.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d7155e39a42102af67fb2bfa7f2c8a6d8fc09e4153e14139000c1b40a87e3dc8
BLAKE2b-256 checksum
How to use checksums
089a80bbd9451f407f579c7117beee48773ac317b6366526c8822a2a00af8a2c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

This release

0.5.1 This release

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page