Skip to main content

Sharada

Sharada

Typed decisions about text in one forward pass — options in the request, calibrated probabilities out.

Apache 2.0 sharada-base on the Hub the write-up

A small encoder that makes typed decisions about text in one forward pass: the options come with the request, and what comes back is a probability for each of them. Nothing is generated, nothing is parsed, and the same model answers a question it has never seen because the labels are part of the input rather than part of the weights.

from sharada import DecisionModel

model = DecisionModel.from_pretrained("lenabarretta/sharada-base")

d = model.decide(
    text="My card still hasn't arrived and I ordered it two weeks ago.",
    question="Which team should handle this?",
    options=["billing", "card delivery", "technical support", "account closure"],
)

d.answer          # 'card delivery'
d.confidence      # 0.86
d.probabilities   # {'billing': 0.07, 'card delivery': 0.86, ...}

Three kinds of question are typed, so the model knows what the options mean:

kind the options are example
choice unordered labels which team, which topic, which intent
scale ordered steps severity 1–5, how positive, how urgent
binary yes or no is this spam, does this need a human
from sharada import Request

model.decide_many([
    Request(text=review, question="How positive is this review?",
            options=["very negative", "negative", "neutral", "positive", "very positive"],
            kind="scale", task="review-tone"),
    Request(text=review, question="Does this mention a refund?",
            options=["no", "yes"], kind="binary", task="refund-flag"),
])

Install

pip install git+https://github.com/LenaBarretta/sharada

Python 3.10+, torch and transformers; CPU is enough to run it.

Models

model parameters download accuracy ECE one decision
lenabarretta/sharada-base 150M 300 MB 0.813 0.008 19.8 ms the default; fine-tune this one
lenabarretta/sharada-large 397M 794 MB 0.834 0.009 26.6 ms better, and better still on label sets it has never seen
lenabarretta/sharada-multilingual-base 308M 616 MB 0.802 0.009 23.1 ms the same model over many more languages, on mmBERT
lenabarretta/sharada-multilingual-small 141M 282 MB 0.773 0.013 21.8 ms half the download, three points behind, and barely faster — for memory, not for latency

Weights are stored in half precision: the encoder was trained under a float16 autocast, so the bits below that were never signal, and the file halves for a shift in the probabilities of about 2e-4. Loading casts back to float32, so fine-tuning is unaffected — this is a storage format, not quantisation, and nothing about the model you get is approximate.

Accuracy is over 38 label sets with every label offered at once — banking on all 77 intents, clinc on all 151 — and ECE is how far the stated probability lands from how often it turns out right. The gap between the two models is widest where it matters most: on label sets neither was trained on, large reads arXiv categories at 0.456 against base's 0.317.

A checkpoint carries its own encoder, limits and temperatures in config.json, so a bigger model — or a multilingual one, built on a multilingual encoder — is another repository rather than another version of the library. It also means a multilingual model cannot be a flag on an English one: this encoder is English down to its tokenizer, and another language means other weights.

Accuracy and calibration per label set are in each model card, measured on label sets the model was not trained on as well as on the ones it was.

What it will not do

  • It does not generate. No free-form answers, no extraction, no reasoning out loud. Options or nothing.
  • 256 tokens of text by default (Limits), with 48 for the question and 12 per option. Longer documents need chunking or a larger limit, and the limit costs quadratic attention.
  • One of them is English-only. base and large are built on an English encoder and were trained on English label sets; the multilingual checkpoints cover far more, but pay about two and a half points of English accuracy for it, and no checkpoint has been measured on a language outside the forty-odd in the mix.
  • No medicine, no law, no code. Nothing of the sort is in the training mix. The two medical label sets are measured only, and one of them — verifying public-health claims — lands at the majority-class baseline, which is to say it does not work. Those domains need fine-tuning on your own labelled examples.
  • It is small. Where a frontier model knows a fact that is not in the text, it wins. This answers questions about the text in front of it.

Seen and unseen

The two numbers worth separating. A label set the model was trained on is the ordinary case; one it has never seen is the claim that makes this architecture interesting at all, since the labels live in the input rather than in the weights.

label sets measured on accuracy calibration error
base — trained on 32 10,432 0.837 0.040
base — never trained on 6 1,056 0.579 0.096
large — trained on 32 10,432 0.855 0.036
large — never trained on 6 1,056 0.626 0.069
multilingual-base — trained on 35 13,960 0.821 0.038
multilingual-base — never trained on 7 1,296 0.600 0.087
multilingual-small — trained on 35 13,960 0.796 0.039
multilingual-small — never trained on 7 1,296 0.519 0.095

Both are weighted by how many examples each label set contributed, so they are the same kind of number as the overall accuracy. The gap is the honest cost of asking a question nobody trained for — and the calibration error roughly doubles across it, which is the part to watch: the model is not only less right on unfamiliar label sets, it is also less aware of being wrong.

Three things the architecture guarantees

The request is laid out as one sequence — text, question, then every option as a parallel branch — and a mask decides who may read whom. Both are in layout.py and masking.py, and they buy three properties that hold by construction, not because training got them approximately right:

  1. The order of the options cannot matter. Every option branch starts at the same position id, and the encoder's positions are rotary, so no option is earlier or later than another. Reshuffle them and the probabilities follow their options exactly.
  2. An option's score does not depend on which other options are offered. An option reads the text, the question and itself, nothing else. Drop two options from a list of four and the remaining two keep their scores to the last bit — so the probabilities are a renormalisation, and a long list of options does not make each one noisier.
  3. The text is read once. Text tokens read only text tokens, so their states do not depend on the question. Ten questions about one document are ten cheap read-outs over one encoding of it.

These are the tests in tests/test_model.py, checked on an untrained model.

Fine-tune it on your own labels

This is what the package is built around: a few hundred labelled examples and a few minutes.

from sharada import DecisionModel, Example

examples = [
    Example(text="the invoice is wrong again", question="Which team should handle this?",
            options=["billing", "technical", "sales"], label=0, task="routing"),
    ...
]

model = DecisionModel.from_pretrained("lenabarretta/sharada-base")
report = model.fit(examples)        # holds out 20%, stops when held-out log loss stops improving
model.save("my-router")

report        # Report(820 examples, accuracy 0.914, log loss 0.287, calibration error 0.031)

fit keeps a part of the examples out, trains on the rest, early-stops, then fits one temperature per task on the held-out part and writes a calibration passport. Useful arguments:

argument default
loss "cross_entropy" "brier" scores the whole distribution, not just the right option
freeze_encoder False train the read-out only: seconds, and enough for a few hundred examples
batch_size, lr, max_epochs, patience 16, 2e-5, 10, 2

Mixing several tasks in one fit is the normal case — give each one its own task name and each gets its own temperature. See examples/finetune_your_own.py, which trains on a CSV.

The number next to the answer is supposed to be true

0.86 should mean right about 86% of the time. That is a property of a distribution, not of a model, so it is fitted and it expires:

from sharada import check_passport, save_passport

passport = model.calibrate(recent_examples)     # one temperature per task
save_passport(passport, "passport.json")

check_passport(passport, recent_examples, model)
# ['the calibration expired on 2026-04-01; fit it again on recent answers',
#  'routing: the options have changed since the calibration']

check_passport returns an empty list when it finds nothing wrong — run it in the deployment pipeline and fail the build on anything it returns. evaluate gives accuracy, log loss, Brier, expected calibration error with a bootstrap interval, a reliability curve and a risk–coverage curve.

What to do with 0.86

A probability is not a decision. Policy turns one into the action that costs the least, including handing the request to a person:

from sharada import Policy, escalation_budget

policy = Policy(options=["approve", "reject"],
                costs={("fraud", "approve"): 10_000, ("clean", "reject"): 100},
                escalate=30)

policy.act({"fraud": 0.1, "clean": 0.9})     # Action('answer', 'reject', expected_loss=90.0)
policy.act({"fraud": 0.5, "clean": 0.5})     # Action('escalate', ...)

# a person can look at 5% of the traffic, no more:
policy = escalation_budget(policy, probabilities, budget=0.05)

Every label set, one by one

Rows marked unseen were kept out of training entirely. A dash means the label set is not in that model's mix. Trained on is what the set contributed to training, measured on how many held-out examples the accuracy rests on — read the second one first, because a row measured on ninety-six examples moves by a full point when a single answer changes.

label set answer options measured on base large multilingual-base multilingual-small trained on what it is
clinc-intent 151 480 0.908 0.950 0.883 0.844 24,000 clinc/clinc_oos/plus
banking-intent 77 480 0.879 0.898 0.854 0.835 19,986 mteb/banking77
massive-intent 59 480 0.867 0.879 0.875 0.865 23,028 mteb/amazon_massive_intent/en
question-type-fine 50 240 0.912 0.904 0.900 0.883 10,904 CogComp/trec
massive-intent-multi 35 1,440 — — 0.849 0.834 72,000 mteb/amazon_massive_intent ×18 languages
fine-emotion 28 360 0.583 0.575 0.581 0.603 18,000 google-research-datasets/go_emotions/simplified
newsgroup 20 292 0.695 0.726 0.685 0.644 14,592 SetFit/20_newsgroups
entity-type 14 360 0.992 0.997 0.997 0.997 18,000 fancyzhx/dbpedia_14
forum-topic 10 360 0.756 0.789 0.769 0.728 18,000 community-datasets/yahoo_answers_topics
topic-multi 7 660 — — 0.815 0.756 30,398 mteb/sib200 ×24 languages
question-type 6 240 0.979 0.963 0.975 0.967 10,904 CogComp/trec
emotion 6 300 0.897 0.910 0.913 0.887 15,000 dair-ai/emotion
sentence-tone 5 300 0.617 0.637 0.570 0.550 15,000 SetFit/sst5
review-stars 5 480 0.646 0.658 0.650 0.640 24,000 Yelp/yelp_review_full
app-stars 5 300 0.720 0.717 0.677 0.687 15,000 sealuzh/app_reviews
tweet-emotion 4 240 0.846 0.879 0.817 0.808 6,514 cardiffnlp/tweet_eval/emotion
news-section 4 360 0.942 0.931 0.933 0.928 18,000 fancyzhx/ag_news
tweet-sentiment 3 300 0.693 0.713 0.680 0.673 15,000 cardiffnlp/tweet_eval/sentiment
entailment-short 3 360 0.892 0.906 0.875 0.850 18,000 stanfordnlp/snli
entailment-multi 3 1,428 — — 0.764 0.691 71,968 facebook/xnli ×14 languages
entailment 3 480 0.833 0.873 0.815 0.779 24,000 nyu-mll/glue/mnli
toxic-comment 2 300 0.930 0.933 0.920 0.923 15,000 SetFit/toxic_conversations
spam 2 240 0.992 0.992 0.992 0.992 10,034 ucirvine/sms_spam
short-verdict 2 240 0.929 0.950 0.938 0.892 12,000 cornell-movie-review-data/rotten_tomatoes
same-question 2 360 0.831 0.867 0.825 0.811 18,000 nyu-mll/glue/qqp
same-meaning 2 240 0.817 0.808 0.787 0.792 7,336 nyu-mll/glue/mrpc
product-tone 2 300 0.957 0.960 0.947 0.927 15,000 fancyzhx/amazon_polarity
paraphrase 2 300 0.907 0.950 0.883 0.883 15,000 google-research-datasets/paws/labeled_final
offensive 2 300 0.847 0.860 0.863 0.863 15,000 cardiffnlp/tweet_eval/offensive
movie-verdict 2 300 0.930 0.957 0.933 0.920 15,000 stanfordnlp/imdb
irony 2 240 0.733 0.787 0.725 0.704 5,724 cardiffnlp/tweet_eval/irony
hateful 2 300 0.833 0.837 0.810 0.817 14,989 cardiffnlp/tweet_eval/hate
grammatical 2 300 0.790 0.803 0.740 0.700 15,000 nyu-mll/glue/cola
follows 2 240 0.817 0.875 0.779 0.696 4,980 nyu-mll/glue/rte
answers-question 2 360 0.894 0.942 0.892 0.853 18,000 nyu-mll/glue/qnli
massive-scenario (unseen) 18 240 0.733 0.742 0.767 0.696 — mteb/amazon_massive_scenario/en
arxiv-category (unseen) 11 180 0.317 0.456 0.522 0.339 — ccdv/arxiv-classification/no_ref
topic-unseen-languages (unseen) 7 240 — — 0.679 0.596 — mteb/sib200 ×8 languages
poem-tone (unseen) 4 96 0.375 0.573 0.510 0.521 — google-research-datasets/poem_sentiment
claim-veracity (unseen) 4 180 0.544 0.467 0.383 0.350 — ImperialCollegeLondon/health_fact
subjective (unseen) 2 180 0.617 0.683 0.461 0.383 — SetFit/subj
medical-pair (unseen) 2 180 0.739 0.772 0.750 0.667 — curaihealth/medical_questions_pairs
overall 15,256 0.813 0.834 0.802 0.773 663,357

Per-label-set calibration error and the fitted temperatures are in each model card.

Train any of them yourself

training/ has the whole run. The mix of label sets is declared in sources.py — intents, topics, review scores, emotion, toxicity, spam, entailment — each one asked through several wordings of its question, with the options shuffled and long label sets often shown as a sampled handful, so the model learns to read the options rather than their positions. Some label sets are kept out of training entirely and only measured: those are the zero-shot numbers above.

Every published checkpoint comes from one of these four commands, with the settings they share written out in full so that nothing about them has to be guessed:

SHARED="--steps 15000 --cap-scale 3 --variants 2 --max-hours 11 \
        --eval-every 1000 --eval-examples 300 --checkpoint-every 500 --keep best"

python training/run.py --encoder answerdotai/ModernBERT-base  --out runs/base  \
    --batch-size 16 --accumulate 2 --lr 3e-5 $SHARED
python training/run.py --encoder answerdotai/ModernBERT-large --out runs/large \
    --batch-size 8 --accumulate 4 --lr 2e-5 $SHARED
python training/run.py --encoder jhu-clsp/mmBERT-base  --out runs/multilingual-base  --multilingual \
    --batch-size 8 --accumulate 4 --lr 3e-5 $SHARED
python training/run.py --encoder jhu-clsp/mmBERT-small --out runs/multilingual-small --multilingual \
    --batch-size 16 --accumulate 2 --lr 3e-5 $SHARED

The effective batch is 32 in all four; the micro-batch differs only because a wider model and a larger vocabulary need more memory to hold. --multilingual adds the label sets that come in many languages and keeps the English ones, so a multilingual model gains languages rather than trading English for them; without the flag they are skipped, which is what keeps the English checkpoints reproducible.

Measured on one free T4: about three hours for base, four to five for large, five for multilingual-base. The multilingual runs also spend twenty minutes assembling the mix before the first step, because it downloads sixty-four language editions.

A multilingual encoder costs more to train than its body suggests. mmBERT-base has exactly the same twenty-two layers and hidden size as ModernBERT-base, so a forward pass is the same work — but the optimiser updates every parameter on every step, and 197M of its 287M are the embedding table. That is twice the optimiser traffic for the same arithmetic, which is where the extra two hours go. At inference the table costs nothing: looking a row up is not a matrix multiply.

The run checkpoints every five hundred steps and resumes from a checkpoint if it finds one, which is what lets a Kaggle session that ran out of time be restarted rather than lost; training/kaggle.ipynb is the notebook wrapper, and a word in its first cell picks which of the four to train. What resuming cannot do is extend a run that already finished: the learning-rate schedule is laid out over the total number of steps, so asking for more lifts it back up, and training/README.md has what happened when that was tried.

Where it came from

The design, the experiments behind it and what each training signal did are written up in RLCR from Scratch, with a runnable lab.

Apache 2.0.

Metadata

Release files for sharada 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sharada 0.1.1
File Size Uploaded
sharada-0.1.1.tar.gz 257.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sharada 0.1.1
File Interpreter ABI Platform
sharada-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 287.2 kB

Release files / sharada-0.1.1.tar.gz

Download URL sharada-0.1.1.tar.gz
Size 257.7 kB
Tags Source
SHA-256 checksum
How to use checksums
5e9df1023a4470edbe7e30522297a0adbdacbd1d8f7ebe967a8de662cc897a3d
BLAKE2b-256 checksum
How to use checksums
2824421e477afaface76ffd9f042b3d9485426210f0d498f8704b6d28cac01ef
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release files / sharada-0.1.1-py3-none-any.whl

Download URL sharada-0.1.1-py3-none-any.whl
Size 29.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
695ead719ebb67ba58a6b73e91a179b0b9895bf5d0c2dd1790635b81abfbc493
BLAKE2b-256 checksum
How to use checksums
81381ca79d5012054b4d1a635abb98c45f9bec674d74ee730fccaab16d567c1d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page