Skip to main content

BaseDecision

Ask a question about a text. Get the answer and how sure the model is. Pick one option, answer yes/no, or give a rating, for texts up to 8,192 tokens. It runs on your own laptop or GPU, with no API key, or through OpenAI/Anthropic. It is a classifier, not a chatbot: it never writes text, it chooses among the answers you give it.

Install

pip install "basedecision[runtime]"    # needs Python 3.10+

That is all for local use (the first-time download of Torch is large).

I want to... Install
Run the model on my machine (CPU or NVIDIA GPU; Apple-silicon Macs use the CPU) pip install "basedecision[runtime]"
...and download the model from Hugging Face pip install "basedecision[runtime,hub]"
Use OpenAI or Anthropic only, no Torch pip install "basedecision[providers]"

Already have Torch 2.6+ and Transformers 4.48-4.57 and want pip to leave them alone? pip install --no-deps basedecision installs only BaseDecision. Nothing is ever installed behind your back: a missing package gives an error that says which extra to install. Working from a clone of the repository instead? Run pip install ".[runtime]" inside it.

Try it

from basedecision import load

model = load('/path/to/model')   # the folder with model.safetensors; GPU if you have one, else CPU
result = model.choose(context='Please refund my purchase.',
                      question='What does the customer request?',
                      options=['Refund', 'Delivery status', 'Change address'])
print(result.answer)          # Refund
print(result.probabilities)   # Refund is about 0.99999, the other two about 0.000002

Loading takes about 8 s and ~2 GiB of RAM; each answer takes about 0.05 s on a laptop CPU for short texts. No flags are needed.

How it performs

BaseDecision against GLiNER 2.5 base, GLiNER2.5-Decide, Decision 1.0 Kai 0.6B and Laya on eight benchmarks (higher is better):

BaseDecision has the highest average score, 57.6 against 38.6 to 41.6, and is ahead on five of eight benchmarks: BANKING77, FinEntity, ContractNLI, VAST and SGD/SGD-X. It is behind on MuSR and NLI4CT, and all models score about 50 on PhishNChips.

  • Highest average: 57.6, against 38.6 to 41.6 for the other four models (the unweighted mean of the eight benchmarks).
  • Ahead on five benchmarks: BANKING77, FinEntity, ContractNLI, VAST and SGD/SGD-X, by 1.2 to 30.2 points. The largest gaps are on ContractNLI and VAST.
  • Behind on two: MuSR (44.6 against 45.2) and NLI4CT (42.3 against 48.0).
  • PhishNChips does not separate the models: all five score between 49.9 and 50.4.

Treat this as a guide, not a guarantee. How a decision model does depends on the task and the wording of its questions, so try it on a sample of your own before relying on it. The probabilities it returns are raw softmax values, not calibrated confidence (see Calibration).

Examples

These examples reuse the model from above.

Yes or no: check

result = model.check(context='The account is active.', question='Is the account active?')
print(result.answer)   # True (a real Python bool)

A rating: score

result = model.score(context='I am satisfied.', question='How satisfied is the customer?',
                     levels=['Dissatisfied', 'Neutral', 'Satisfied'], values=[1, 2, 3])
print(result.answer, round(result.expected_value, 2))   # Satisfied 3.0

Several questions about one text: decide

from basedecision import decide

answers = decide(model, context='I was charged twice. Please refund the duplicate payment.',
                 questions={
                     'intent': {'kind': 'choice', 'question': 'What does the customer request?',
                                'options': ['Refund request', 'Delivery status', 'Other']},
                     'duplicate': {'kind': 'noul', 'question': 'Was the customer charged twice?'},
                 })
print(answers['intent'].answer, answers['duplicate'].answer)   # Refund request True

kind is choice, noul (yes/no) or score. Each question is answered independently.

Options with descriptions: Option

Plain strings are fine. If a label is ambiguous, add a description:

from basedecision import Option

result = model.choose(context='Where is my parcel?', question='What is requested?',
                      options=[Option('refund', 'Customer wants money back'),
                               Option('status', 'Customer asks where the parcel is')])
print(result.answer, '-', result.label)   # status - Customer asks where the parcel is

Many texts at once: predict_batch, predict_iter

from basedecision import Request

options = (Option('refund', 'Refund'), Option('status', 'Delivery status'))
requests = [Request('Please refund this purchase.', 'What is requested?', options),
            Request('Where is my parcel?', 'What is requested?', options)]
print([r.answer for r in model.predict_batch(requests)])    # ['refund', 'status']
for result in model.predict_iter(requests):                 # same, one at a time, for long streams
    print(result.answer)

Everything in a result

print(model.choose(context='Please refund my purchase.', question='What is requested?',
                   options=['Refund', 'Other']).to_dict())
# answer, option_id, label, probabilities, raw_logits, packed_tokens, kind, ...  (plain JSON types)

probabilities are the model's raw softmax, not calibrated.

Download the model from Hugging Face: load_from_hub

Needs pip install "basedecision[runtime,hub]". The repository name is a placeholder until the model is published.

from basedecision import load_from_hub

model = load_from_hub('YOUR_NAMESPACE/YOUR_MODEL', revision='YOUR_COMMIT_HASH')

A commit hash pins the exact weights. Add local_files_only=True to use only the local cache. Sign-in uses the standard Hugging Face setup; no remote Python code is ever run.

Force the CPU or the GPU

model = load('/path/to/model', device='cpu', precision='fp32')
print(model.device, model.precision)    # cpu fp32
# load('/path/to/model', device='cuda')   # force the GPU; an error if there is none

load picks CUDA with BF16 when a suitable NVIDIA GPU exists, else the CPU with FP32. An explicit choice is never second-guessed. Rough CPU timings on one 10-core Apple-silicon laptop (a guide, not a benchmark):

Text length Time per answer Peak memory
short (a few dozen tokens) ~0.05 s ~2 GiB
512 tokens ~0.4 s ~2 GiB
4,096 tokens ~8 s ~5 GiB
8,191 tokens (the limit) ~30 s ~8 GiB

Experimental CPU mode: cpu_fast

from basedecision import load, CPUFastUnavailable

try:
    model = load('/path/to/model', backend='cpu_fast')
except CPUFastUnavailable:                  # unsupported Torch/Transformers: nothing is changed
    model = load('/path/to/model')          # the normal CPU path

Opt-in and experimental: it checks itself at load time and never falls back silently. It was developed on an Intel Xeon server and was 10-15x slower than the default on an Apple-silicon laptop, so measure before you use it. See CPU support and limitations.

OpenAI or Anthropic instead of the local model

export OPENAI_API_KEY=...     # or ANTHROPIC_API_KEY; needs: pip install "basedecision[providers]"
from basedecision import BaseDecision

with BaseDecision.from_provider('openai', 'YOUR_OPENAI_MODEL_ID') as cloud:
    result = cloud.choose(context='Please refund this purchase.',
                          question='What does the customer request?',
                          options=['Refund request', 'Delivery status', 'Other request'])
    print(result.answer)   # same methods as the local model; swap in 'anthropic' for Anthropic

Your text is sent to that provider, and never as an automatic fallback. Cloud results have no probabilities (result.probabilities is None). Reasoning models, errors, retries and costs: docs/CLOUD.md.

Calibrated probabilities (optional, GPU only)

Raw probabilities are the default. For the assessed SGD service-intent workloads only, you can opt in to calibrated ones. This needs a CUDA GPU with BF16 and handles choice questions only.

from basedecision import CalibratedDecision, calibration_profiles

print(list(calibration_profiles()))   # ['clinc', 'sgd_identifier', 'sgd_schema', 'vast']
calibrated = CalibratedDecision(model, profile='sgd_schema')   # raises on a CPU-only machine
result = calibrated.predict(sgd_request)    # a correctly formatted SGD request
print(result.probabilities, result.raw_probabilities)

These profiles do not apply to other questions. Coverage and caveats: docs/CALIBRATION.md.

Answer Jev / SystemOne requests: SystemOne

For the Jev/SystemOne wire format (a state and typed questions):

from basedecision import SystemOne

service = SystemOne(model)
response = service({'model': 'basedecision',
                    'state': 'Our checkout started returning errors and orders are blocked.',
                    'questions': {
                        'department': {'type': 'choice', 'instructions': 'Which team should handle it?',
                                       'criteria': {'billing': 'Payments or invoices',
                                                    'technical': 'Bugs or outages'}},
                        'outage': {'type': 'noul', 'instructions': 'Is a service down?'}}})
print(response['answers']['department']['choice'])   # technical

Or serve it over HTTP with the reference server in the repository's examples/ folder (bound to localhost; see docs/SYSTEMONE.md before exposing it):

python examples/systemone_server.py --model /path/to/model
curl -s localhost:8080/v1/systemone -H 'Content-Type: application/json' \
  -d '{"model":"basedecision","state":"I was charged twice. Please refund me.","questions":{"refund":{"type":"noul","instructions":"Does the customer want a refund?"}}}'

Text only, one forward pass per question, no truncation, and no cloud providers (they return no probabilities).

Free the memory when you are done: close

with load('/path/to/model') as scratch:     # the weights are freed when the block ends
    print(scratch.check(context='The account is active.', question='Is the account active?').answer)   # True
print(scratch.closed)   # True: asking a closed model for an answer raises InputError

model.close() does the same without a with block, and is safe to call twice. It waits for a request that is running, then frees the weights and, on a GPU, the memory they held. The cloud backend and cpu_fast models work the same way.

Limits and errors

from basedecision import ContextLengthError, InputError

print(model.count_tokens(requests[0]))   # 22: how many tokens a request uses, before you run it
try:
    model.choose(context='word ' * 20000, question='Which?', options=['a', 'b'])
except ContextLengthError as error:
    print(error.required_tokens, error.maximum_tokens)   # 20014 8192
except InputError as error:                              # bad options, wrong types, duplicates
    print(error)
  • The limit is 8,192 tokens in total: the question, every option and the text. Nothing is ever silently truncated; an over-long input is a ContextLengthError you can handle, for example by splitting the text.
  • InputError means the request itself is wrong. The context must be a string (serialize structured data yourself), and option ids and labels must be distinct.
  • A cloud call that fails raises ProviderError with code, retryable and status_code, and never includes your key or text (docs/CLOUD.md).
  • A refused or invalid answer is an error, never a guessed label or a silent False.

More

From a clone of the repository, to run the tests:

python -m unittest discover -s tests -v                                          # no model needed
BASEDECISION_TEST_MODEL=/path/to/model python -m unittest discover -s tests -v   # + the real model

With a checkpoint, the tests run the examples above in order and check the output they print. The Hugging Face, cloud and GPU-calibration examples cannot run on a laptop, so they are only checked for syntax and imports.

BaseDecision is a decision model (ModernBERT-large encoder with a typed decision head, 8,192-token window). It has no .generate(), no AutoModelForCausalLM, no chat endpoint, and no reasoning comparable to a general chat model.


License

Apache 2.0. Developed by Hrudayaditya "Aady" Jallu.

Metadata

Release files for basedecision 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for basedecision 0.1.0
File Size Uploaded
basedecision-0.1.0.tar.gz 1.9 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for basedecision 0.1.0
File Interpreter ABI Platform
basedecision-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.0 MB

Release files / basedecision-0.1.0.tar.gz

Download URL basedecision-0.1.0.tar.gz
Size 1.9 MB
Tags Source
SHA-256 checksum
How to use checksums
292791b46738424bf82d18e107dd29e72132c485f64e1ae181653ed890fbb2cb
BLAKE2b-256 checksum
How to use checksums
97aeb9a0be62321e505cb6a87092b3facd44ab26dabc720397757160ce6bf1b3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / basedecision-0.1.0-py3-none-any.whl

Download URL basedecision-0.1.0-py3-none-any.whl
Size 62.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a7ac0c68ac13b511cadb5a2b84b73d56dcee310040e0abfa11904a63dfbca91d
BLAKE2b-256 checksum
How to use checksums
73d6ab62e4e2249f86469986dfd4f76b9465983650921a06e59776363b2a943b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page