Skip to main content

Sharada

Sharada

Typed decisions about text in one forward pass — options in the request, calibrated probabilities out.

Apache 2.0 sharada-base on the Hub the write-up

A small encoder that makes typed decisions about text in one forward pass: the options come with the request, and what comes back is a probability for each of them. Nothing is generated, nothing is parsed, and the same model answers a question it has never seen because the labels are part of the input rather than part of the weights.

from sharada import DecisionModel

model = DecisionModel.from_pretrained("lenabarretta/sharada-base")

d = model.decide(
    text="My card still hasn't arrived and I ordered it two weeks ago.",
    question="Which team should handle this?",
    options=["billing", "card delivery", "technical support", "account closure"],
)

d.answer          # 'card delivery'
d.confidence      # 0.86
d.probabilities   # {'billing': 0.07, 'card delivery': 0.86, ...}

Three kinds of question are typed, so the model knows what the options mean:

kind the options are example
choice unordered labels which team, which topic, which intent
scale ordered steps severity 1–5, how positive, how urgent
binary yes or no is this spam, does this need a human
from sharada import Request

model.decide_many([
    Request(text=review, question="How positive is this review?",
            options=["very negative", "negative", "neutral", "positive", "very positive"],
            kind="scale", task="review-tone"),
    Request(text=review, question="Does this mention a refund?",
            options=["no", "yes"], kind="binary", task="refund-flag"),
])

Install

pip install git+https://github.com/LenaBarretta/sharada

Python 3.10+, torch and transformers; CPU is enough to run it.

Models

model parameters accuracy ECE one decision
lenabarretta/sharada-base 150M 0.813 0.008 19.8 ms the default; fine-tune this one
lenabarretta/sharada-large 397M 0.834 0.009 26.6 ms better, and better still on label sets it has never seen
lenabarretta/sharada-multilingual (in training) 300M the same architecture over 1800+ languages, on mmBERT

Accuracy is over 38 label sets with every label offered at once — banking on all 77 intents, clinc on all 151 — and ECE is how far the stated probability lands from how often it turns out right. The gap between the two models is widest where it matters most: on label sets neither was trained on, large reads arXiv categories at 0.456 against base's 0.317.

A checkpoint carries its own encoder, limits and temperatures in config.json, so a bigger model — or a multilingual one, built on a multilingual encoder — is another repository rather than another version of the library. It also means a multilingual model cannot be a flag on an English one: this encoder is English down to its tokenizer, and another language means other weights.

Accuracy and calibration per label set are in each model card, measured on label sets the model was not trained on as well as on the ones it was.

Three things the architecture guarantees

The request is laid out as one sequence — text, question, then every option as a parallel branch — and a mask decides who may read whom. Both are in layout.py and masking.py, and they buy three properties that hold by construction, not because training got them approximately right:

  1. The order of the options cannot matter. Every option branch starts at the same position id, and the encoder's positions are rotary, so no option is earlier or later than another. Reshuffle them and the probabilities follow their options exactly.
  2. An option's score does not depend on which other options are offered. An option reads the text, the question and itself, nothing else. Drop two options from a list of four and the remaining two keep their scores to the last bit — so the probabilities are a renormalisation, and a long list of options does not make each one noisier.
  3. The text is read once. Text tokens read only text tokens, so their states do not depend on the question. Ten questions about one document are ten cheap read-outs over one encoding of it.

These are the tests in tests/test_model.py, checked on an untrained model.

Fine-tune it on your own labels

This is what the package is built around: a few hundred labelled examples and a few minutes.

from sharada import DecisionModel, Example

examples = [
    Example(text="the invoice is wrong again", question="Which team should handle this?",
            options=["billing", "technical", "sales"], label=0, task="routing"),
    ...
]

model = DecisionModel.from_pretrained("lenabarretta/sharada-base")
report = model.fit(examples)        # holds out 20%, stops when held-out log loss stops improving
model.save("my-router")

report        # Report(820 examples, accuracy 0.914, log loss 0.287, calibration error 0.031)

fit keeps a part of the examples out, trains on the rest, early-stops, then fits one temperature per task on the held-out part and writes a calibration passport. Useful arguments:

argument default
loss "cross_entropy" "brier" scores the whole distribution, not just the right option
freeze_encoder False train the read-out only: seconds, and enough for a few hundred examples
batch_size, lr, max_epochs, patience 16, 2e-5, 10, 2

Mixing several tasks in one fit is the normal case — give each one its own task name and each gets its own temperature. See examples/finetune_your_own.py, which trains on a CSV.

The number next to the answer is supposed to be true

0.86 should mean right about 86% of the time. That is a property of a distribution, not of a model, so it is fitted and it expires:

from sharada import check_passport, save_passport

passport = model.calibrate(recent_examples)     # one temperature per task
save_passport(passport, "passport.json")

check_passport(passport, recent_examples, model)
# ['the calibration expired on 2026-04-01; fit it again on recent answers',
#  'routing: the options have changed since the calibration']

check_passport returns an empty list when it finds nothing wrong — run it in the deployment pipeline and fail the build on anything it returns. evaluate gives accuracy, log loss, Brier, expected calibration error with a bootstrap interval, a reliability curve and a risk–coverage curve.

What to do with 0.86

A probability is not a decision. Policy turns one into the action that costs the least, including handing the request to a person:

from sharada import Policy, escalation_budget

policy = Policy(options=["approve", "reject"],
                costs={("fraud", "approve"): 10_000, ("clean", "reject"): 100},
                escalate=30)

policy.act({"fraud": 0.1, "clean": 0.9})     # Action('answer', 'reject', expected_loss=90.0)
policy.act({"fraud": 0.5, "clean": 0.5})     # Action('escalate', ...)

# a person can look at 5% of the traffic, no more:
policy = escalation_budget(policy, probabilities, budget=0.05)

Train the base models yourself

training/ has the whole run: a mix of public label sets in sources.py — intents, topics, sentiment, toxicity, spam, entailment, review scores — each one presented with several question phrasings, shuffled options and sampled option subsets, so the model learns to read the options rather than their positions. Some label sets are held out of training entirely and only measured, which is where the zero-shot numbers come from.

python training/run.py --encoder answerdotai/ModernBERT-base --out runs/base

It checkpoints every few hundred steps and resumes from the checkpoint if it finds one, which is what makes it survive a Kaggle session; training/kaggle.ipynb is the notebook wrapper. One free T4: about an hour for base, about three for large.

What it will not do

  • It does not generate. No free-form answers, no extraction, no reasoning out loud. Options or nothing.
  • 256 tokens of text by default (Limits), with 48 for the question and 12 per option. Longer documents need chunking or a larger limit, and the limit costs quadratic attention.
  • English. The encoder is English-only; a multilingual encoder drops in, but the published checkpoints are not trained for it.
  • It is small. Where a frontier model knows a fact that is not in the text, it wins. This answers questions about the text in front of it.

Where it came from

The design, the experiments behind it and what each training signal did are written up in RLCR from Scratch, with a runnable lab.

Apache 2.0.

Metadata

Release files for sharada 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sharada 0.1.0
File Size Uploaded
sharada-0.1.0.tar.gz 249.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sharada 0.1.0
File Interpreter ABI Platform
sharada-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 274.9 kB

Release files / sharada-0.1.0.tar.gz

Download URL sharada-0.1.0.tar.gz
Size 249.1 kB
Tags Source
SHA-256 checksum
How to use checksums
7cd351c49b9d8483a61971152a4e9ec2002a629ebd187f86cf8af5a04762f5b5
BLAKE2b-256 checksum
How to use checksums
73369d5d37f48cba3a489e44498b4744c9ef8f04b2875c5a97bd760d9dc2cbc4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / sharada-0.1.0-py3-none-any.whl

Download URL sharada-0.1.0-py3-none-any.whl
Size 25.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cf61b9de4f39e1f9848a5599091c37d122421272d244aff68096ce03bd8c6c87
BLAKE2b-256 checksum
How to use checksums
5c0868024a7eaa556d0313fa7e73e815ffa31f6d896c35cf6e8c85df68a28d5f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page