OpenDecision
An open-weight, local decision model for software.
LLMs return strings. Software needs a branch: which path, how sure, act or escalate. OpenDecision takes arbitrary state — text, JSON, logs, key/value dumps, conversations — plus a typed question, and returns a typed, probabilistic decision with calibrated confidence.
State + Typed Question -> Probabilistic Decision
curl -s localhost:8000/v1/decisions -H 'content-type: application/json' -d '{
"state": {"customer_message": "When will my new card arrive?"},
"question": {"type": "choice", "question": "What is the customer'\''s intent?",
"options": ["card_arrival", "card_not_working", "exchange_rate", "top_up_failed"]}}'
{"type": "choice", "decision": "card_arrival",
"probabilities": {"card_arrival": 0.978, "card_not_working": 0.018,
"exchange_rate": 0.002, "top_up_failed": 0.003},
"confidence": 0.978, "abstained": false, "model": "open-decision-150m"}
It runs entirely on your machine: ~150M parameters, ~10 ms per decision on a GPU, no API key, no telemetry, no data leaving the host.
Why this exists
A general LLM can answer that question, but you pay for it in latency, cost, and — most importantly — you get prose you must parse, with a confidence number that means nothing. For the narrow, repetitive decisions that make up real workflows:
which queue does this ticket go to? -> choice
is this incident customer-impacting? -> binary
should we auto-approve, review, or reject? -> choice + confidence threshold
…a small encoder is the right tool. OpenDecision gives you:
- a schema you can trust — the decision is always one of your declared options, the probabilities are valid and sum to 1, and no free-form text is ever returned;
- calibrated confidence — ECE 0.047, so a 0.9 really is about 90% (see Benchmarks);
- abstention — below your threshold it returns
nullinstead of guessing; - order invariance by construction — shuffling your options changes nothing (measured flip rate: 0.0);
- local-first — after the model is downloaded, no internet is required at all.
OpenDecision explores the same broad problem as recent "machine-native" / "System One"-style decision models, from an open, local-first perspective. It is an independent implementation, not a reproduction of any proprietary system.
Get the model
The weights are not in this repository. Download the archive and point the server at it —
it is fetched once, verified, extracted and cached under ~/.cache/open-decision/.
Model:
open-decision-150mv0.2.0 · ModernBERT-base · Apache-2.0 · 558 MB zip Download:https://drive.google.com/file/d/<FILE_ID>/viewsha256:<paste the checksum here>
Option A — let the server download it (recommended)
open-decision serve --model 'gdrive:<FILE_ID>#sha256=<checksum>'
gdrive:<FILE_ID> works with a bare id or a full Drive share link. Any plain URL works too, so
you can host the archive anywhere:
open-decision serve --model 'https://example.com/open-decision-150m-v0.2.0.zip#sha256=<checksum>'
The #sha256= fragment is optional but recommended — the download is rejected on a mismatch.
Re-running is instant; the cache is reused and the network is never touched again.
Set OPEN_DECISION_CACHE to move the cache, and --offline to forbid any download.
Option B — download it yourself
Unzip it anywhere and pass the directory:
unzip open-decision-150m-v0.2.0.zip -d ./models
open-decision serve --model ./models/model
The directory must contain config.json, model.safetensors, open_decision.json and the
tokenizer files. Only safetensors weights are loaded and trust_remote_code is never
enabled, so nothing in a model directory can execute code.
Install
Heads-up before you run anything:
pip install open-decision/uv tool install open-decisionwill not work yet — the package is not on PyPI. And a default install pulls the CUDA build of PyTorch: ~5.4 GB, which downloads for minutes with almost no output and looks like a freeze. Every command below pins the CPU build (~950 MB, ~1 minute) unless you actually need a GPU.
Step 1 — install uv
curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
# Windows: powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
Restart the shell, then check: uv --version (0.8 or newer).
Step 2 — get the code
Until the package is published, install from a checkout:
git clone https://github.com/<gh-user>/open-decision.git
cd open-decision
Step 3 — install
Pick one.
A. As a command-line tool (a global open-decision binary, isolated environment):
uv tool install --index pytorch-cpu=https://download.pytorch.org/whl/cpu \
--index-strategy unsafe-best-match .
Note uv tool install ignores UV_TORCH_BACKEND, so the --index flags above are how you get
the small CPU build. Without them you get the 5.4 GB CUDA build.
B. As a project (for development, or to use it as a library):
uv sync --torch-backend cpu # ~950 MB, about a minute
uv run open-decision --version
C. In your own virtualenv:
uv venv && uv pip install --torch-backend cpu .
With a GPU — drop the CPU pins and let uv detect your CUDA version. Expect a multi-GB
download; run it with -v so you can see progress:
uv sync --torch-backend auto -v
Verify whichever you chose:
open-decision --version # or: uv run open-decision --version
Step 4 — get the model and serve it
See Get the model above for the download link and checksum.
open-decision serve --model 'gdrive:<FILE_ID>#sha256=<checksum>' --abstain-threshold 0.7
First start downloads ~558 MB and prints progress; later starts are instant from the cache.
Wait for {"status":"ready"} on /ready, then:
curl -s localhost:8000/v1/decisions -H 'content-type: application/json' -d '{
"state": {"customer_message": "My card is not working at the ATM"},
"question": {"type": "choice", "question": "What is the customer'"'"'s intent?",
"options": ["card_arrival", "card_not_working", "exchange_rate"]}}'
Troubleshooting installs
| symptom | cause and fix |
|---|---|
uv tool install open-decision hangs, then fails |
not on PyPI yet — install from a checkout (Step 3) |
| install seems frozen for minutes | downloading the CUDA PyTorch build (~5.4 GB). Ctrl-C and add the CPU pins above, or add -v to watch progress |
open-decision: command not found |
~/.local/bin is not on PATH; run uv tool update-shell and restart the shell |
| server returns HTML error on model download | the Drive file is not shared as "Anyone with the link", or the daily quota is exhausted — download it manually and pass the directory |
/ready returns 503 for a long time |
still downloading or loading the model; watch the server log |
| slow decisions on CPU | expected — CPU is roughly 10x slower than GPU. Use --device cuda with a CUDA install, or Docker |
No-install option: Docker
Nothing to install but Docker itself — the model is fetched on first start and cached in the mounted volume:
# GPU
docker run --gpus all -p 8000:8000 -v ~/.cache/open-decision:/root/.cache/open-decision \
-e OPEN_DECISION_MODEL='gdrive:<FILE_ID>' ghcr.io/<gh-user>/open-decision:latest
# CPU
docker run -p 8000:8000 -v ~/.cache/open-decision:/root/.cache/open-decision \
-e OPEN_DECISION_MODEL='gdrive:<FILE_ID>' ghcr.io/<gh-user>/open-decision:cpu
Using it
Python SDK
from open_decision import Client
client = Client("http://localhost:8000")
result = client.decide(
state={"customer_message": "My card is not working at the ATM"},
question={"type": "choice", "question": "What is the customer's intent?",
"options": ["card_arrival", "card_not_working", "exchange_rate"]},
)
print(result.decision, result.confidence, result.probabilities)
AsyncClient has the same methods; concurrent calls are batched into one forward pass.
In-process, no server
from open_decision import DecisionEngine
engine = DecisionEngine.load("gdrive:<FILE_ID>") # or a local directory
decision = engine.decide(
state="status=failed retry_count=4 region=eu-west",
question={"type": "binary", "question": "Should this job be retried again?"},
)
print(decision.decision, decision.probability)
Several decisions about one state
Workflows usually need more than one narrow decision. They are all batched into a single forward pass:
{"state": {"...": "..."},
"questions": [
{"id": "intent", "type": "choice", "options": ["refund", "reship", "investigate"]},
{"id": "urgent", "type": "binary", "question": "Is this urgent?"}
],
"abstain_threshold": 0.7}
{"decisions": {"intent": {"...": "..."}, "urgent": {"...": "..."}},
"model": "open-decision-150m"}
Acting on confidence
The whole point is that you can branch on the number:
d = client.decide(state=ticket, question={"type": "choice", "options": [...]})
if d.abstained or d.confidence < 0.60:
route_to_human(ticket)
elif d.confidence < 0.80:
queue_for_review(d.decision, ticket)
else:
execute(d.decision, ticket)
Benchmarks
v0.2.0, on a held-out test set of 4,660 examples (synthetic + Banking77 + CLINC150 + TREC).
Reproduce with the training repo's evaluate command; full numbers in runs/v0.2/eval.json.
Against baselines
| accuracy | ECE ↓ | NLL ↓ | |
|---|---|---|---|
| OpenDecision 150M | 0.670 | 0.047 | 0.718 |
| TF-IDF + logistic regression | 0.367 | 0.095 | 1.423 |
| majority class | 0.350 | 0.105 | 1.515 |
+30 accuracy points over a bag-of-words cross-encoder on the same data.
By dataset
| dataset | n | accuracy | ECE | notes |
|---|---|---|---|---|
| CLINC150 | 562 | 0.932 | 0.022 | 150-way intent (8 options sampled) |
| Banking77 | 596 | 0.923 | 0.027 | 77-way banking intent |
| synthetic | 1712 | 0.685 | 0.102 | 20 domains, 3 languages |
| TREC | 451 | 0.333 | 0.055 | zero-shot — never trained on (chance 0.167) |
The headline 0.670 blends supervised intent classification with zero-shot transfer. For in-domain intent classification the model is in the low 90s.
Confidence is the point
Selective accuracy — the model answers only when confident enough:
| threshold | coverage | accuracy on answered |
|---|---|---|
| 0.6 | 50.5% | 91.5% |
| 0.7 | 41.5% | 95.3% |
| 0.8 | 34.9% | 97.2% |
| 0.9 | 28.2% | 98.2% |
Confidence separates answerable from unanswerable input: feeding it garbage
("asdf qwer", {"a": 1}, "ok") yields 0.29–0.59 confidence versus ~1.00 on a clear intent
— an AUC of 1.00 on banking intents and 0.84 on open-ended routing questions.
Other properties
| property | value |
|---|---|
| option-order invariance | flip rate 0.0, mean JSD 3e-16 (500 examples) |
| held-out domains | 0.600 vs 0.702 on seen domains |
| languages | en 0.698 · fr 0.609 · ar 0.539 |
| state formats | text 0.716 · logs 0.668 · kv 0.628 · conversation 0.560 · json 0.517 |
| throughput | 326 decisions/s (RTX 5090, bf16) |
| memory | 0.96 GB GPU at inference |
Suggested policy
confidence >= 0.80 -> execute automatically (97.2% accurate, ~35% of traffic)
0.60 - 0.80 -> execute with review
confidence < 0.60 -> route to a human
Calibrate these on your own labelled sample — the numbers above come from this test set.
Known limitations — please read
scorequestions are not functional in v0.2. The score head is effectively a constant predictor (0.47 spread on a 1–10 scale). The API accepts them; do not rely on the output.binaryis weaker thanchoice— 0.638 accuracy and ECE 0.136 (choice: 0.045). Usable with a high threshold and review, not for unattended automation.- Arabic (0.539) lags English (0.698); JSON and conversation-shaped states are the weakest formats.
- Confidence is not comparable across question types. A "real" banking intent scores ~1.00 while an open-ended routing question scores ~0.63. Pick the threshold per workflow.
- Trained largely on synthetic data distilled from one open-weight teacher, so it inherits that teacher's biases. Validate on your own data before trusting it.
Decision types
| type | request | response |
|---|---|---|
choice |
{"type":"choice","options":["A","B","C"],"question":"optional"} |
decision (one of the options), probabilities, confidence |
binary (alias noul) |
{"type":"binary","question":"Should this be escalated?"} |
decision (bool), probability (P(yes)), confidence |
score |
{"type":"score","question":"How urgent?","min":1,"max":10} |
score, std, confidence — not functional in v0.2 |
Below abstain_threshold the decision is null and abstained is true; the probabilities
are still returned so your application can apply its own policy.
API
| Method | Path | Purpose |
|---|---|---|
POST |
/v1/decisions |
decisions for one state and one or more questions |
GET |
/v1/models |
model id, version, device, dtype, calibration |
GET |
/health |
process is up |
GET |
/ready |
model loaded (503 while loading / downloading) |
GET |
/docs, /openapi.json |
interactive docs and OpenAPI schema |
Full schema in docs/api.md; runnable samples in examples/.
CLI
open-decision serve --model <dir | gdrive:ID | https://...> [--port 8000] [--abstain-threshold 0.7]
open-decision health
open-decision model-info
echo '{"state":"...","question":{"type":"binary","question":"..."}}' | open-decision decide
echo '{...}' | open-decision decide --model ./models/model # no server needed
Architecture
┌──────────────────── OpenDecision model ─────────────────────┐
state ─────┐ │ per option: "[CLS] question / option [SEP] state [SEP]" │
question ──┼──▶ │ ModernBERT encoder ──▶ [CLS] ──▶ choice head ──▶ logit│──▶ softmax
options ───┘ └────────────────────────────────────────────────────────────┘ over options
│
temperature scaling (optional) ──▶ abstention threshold ──▶ typed decision
Choice and binary questions are scored one option at a time by a cross-encoder, which is why the model is invariant to option order by construction and accepts any option strings rather than a fixed label set. See docs/model-format.md.
Local-first / privacy
- Inference runs entirely on your machine. After the model is cached,
--offlineguarantees no network access. - No telemetry, no external API calls, no request-body logging by default.
- Optional API key, CORS allow-list, body-size limit, max state length and question count.
Details in docs/privacy.md.
Configuration
Every setting is a CLI flag or an OPEN_DECISION_* environment variable:
| Setting | Default | Notes |
|---|---|---|
OPEN_DECISION_MODEL |
– | directory, Hub id, gdrive:<id>, or archive URL (required) |
OPEN_DECISION_CACHE |
~/.cache/open-decision |
where downloaded models are stored |
OPEN_DECISION_HOST / PORT |
127.0.0.1 / 8000 |
Docker images use 0.0.0.0 |
OPEN_DECISION_DEVICE / DTYPE |
auto |
cpu, cuda, mps |
OPEN_DECISION_ABSTAIN_THRESHOLD |
0.0 |
never abstain by default |
OPEN_DECISION_CALIBRATION |
true |
apply stored temperature scaling |
OPEN_DECISION_BATCHING |
true |
dynamic batching across requests |
OPEN_DECISION_API_KEY |
– | enables Authorization: Bearer / X-API-Key |
OPEN_DECISION_OFFLINE |
false |
never download anything |
OPEN_DECISION_LOG_REQUESTS |
false |
request bodies are never logged unless enabled |
Development
uv sync
uv run pytest -q
uv run ruff check src tests && uv run mypy src
Tests build a tiny random model on the fly; no downloads required.
Roadmap
- fix the score head; strengthen binary
- better Arabic and structured-state (JSON) performance
- larger encoder option, quantization (int8/fp8), ONNX
- multi-label and ranking decision types
License
Apache-2.0 for the code. Model weights are released under Apache-2.0 and carry their own model card; third-party datasets used in training keep their own licenses.
Release files for open-decision-ai 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| open_decision_ai-0.2.0.tar.gz | 164.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| open_decision_ai-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 201.6 kB
Release files / open_decision_ai-0.2.0.tar.gz
| Download URL | open_decision_ai-0.2.0.tar.gz |
|---|---|
| Size | 164.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
54dbb05635d62d63cc2259a02224acbd51401a31bffd80a4ce84c2d8ba698d4c
|
|
BLAKE2b-256 checksum How to use checksums |
cc305714f9dd420d8ec62b0825736a46e542d1e9a059760320d88c0d4b99b794
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.12
|
Release files / open_decision_ai-0.2.0-py3-none-any.whl
| Download URL | open_decision_ai-0.2.0-py3-none-any.whl |
|---|---|
| Size | 37.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3ad5f12607b9f05cd21d37c03b6576b4e135de31a0c21310c2edac73e19c679f
|
|
BLAKE2b-256 checksum How to use checksums |
c24e4e33ea01dd427df81e45c575da32c81663aa37781820eb1f2a4c7c83d97d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.8.12
|