lenabarretta/sharada-large

A small encoder that makes typed decisions about text in one forward pass. The options come with the request and what comes back is a probability for each of them: nothing is generated and nothing is parsed. Because the labels are part of the input rather than part of the weights, a label set it has never been trained on still gets an answer.

Code, examples and the training run: https://github.com/LenaBarretta/sharada. The design and the experiments behind it: https://lenatriestounderstand.com/notes/llm/024-rlcr/.

pip install sharada
from sharada import DecisionModel

model = DecisionModel.from_pretrained("lenabarretta/sharada-large")

d = model.decide(
    text="My card still hasn't arrived and I ordered it two weeks ago.",
    question="Which team should handle this?",
    options=["billing", "card delivery", "technical support", "account closure"],
)
d.answer, d.confidence, d.probabilities

Fine-tuning on a few hundred of your own labelled examples is the intended use:

report = model.fit(examples)      # holds 20% out, early-stops, then fits a temperature per task
model.save("my-router")

What it is

encoder answerdotai/ModernBERT-large
parameters 397M
text up to 256 tokens, plus 48 for the question and 12 per option
kinds of question choice (unordered labels), scale (ordered steps), binary (yes or no)
output one probability per option, from a single forward pass
weights float16 on disk, float32 in memory — a storage format, not quantisation
one decision 26.6 ms per question

Three properties hold by construction rather than by training: the order of the options cannot change the answer (every option branch starts at the same position id), an option's score does not depend on which other options are offered (an option reads the text, the question and itself), and the text is read once however many questions are asked of it. The tests in the repository check all three on an untrained model.

How it was trained

488,991 examples from 32 public label sets — intents, topics, review scores, emotion, toxicity, spam and entailment — for 15,000 steps of 32, AdamW at 2e-05 with a cosine schedule, cross entropy loss. The published weights are the ones that measured best on held-out data, at step 15,000 of 15,000 — past that the model stops answering better and only grows more certain. 6 further label sets were held out of training entirely and only measured: arxiv-category, claim-veracity, massive-scenario, medical-pair, poem-tone, subjective.

In training the options were shuffled, long label sets were often shown as a sampled handful, and each label set was asked through several wordings of its question — so the model reads the options and the question rather than their positions.

What it scores

Trained on is what that label set actually contributed to this run — not its cap, since a small dataset runs out first and a pooled multilingual one is split between its languages; a dash means the model never saw it. Measured on is how many held-out examples the accuracy beside it rests on, and it is worth reading first: a row measured on ninety-six examples moves by a full point when one answer changes.

Measured on held-out examples, with every label offered at once — all 151 intents of clinc, all 77 of banking — because that is what a request actually looks like. One temperature per label set was fitted on the same held-out examples. ECE is the expected calibration error over 15 equal bands: how far the stated probability is from how often it turns out right. Rows marked unseen are label sets kept out of training entirely, never trained on, only measured.

label set answer options kind trained on measured on accuracy log loss ECE T
clinc-intent 151 choice 24,000 480 0.950 0.226 0.022 0.891
banking-intent 77 choice 19,986 480 0.898 0.380 0.019 1.122
massive-intent 59 choice 23,028 480 0.879 0.418 0.033 1.26
question-type-fine 50 choice 10,904 240 0.904 0.381 0.042 1.587
fine-emotion 28 choice 18,000 360 0.575 1.235 0.062 1.122
newsgroup 20 choice 14,592 292 0.726 0.798 0.052 1.26
entity-type 14 choice 18,000 360 0.997 0.016 0.005 0.891
forum-topic 10 choice 18,000 360 0.789 0.677 0.058 1.0
question-type 6 choice 10,904 240 0.963 0.110 0.015 1.414
emotion 6 choice 15,000 300 0.910 0.189 0.034 1.335
app-stars 5 scale 15,000 300 0.717 0.791 0.049 1.0
review-stars 5 scale 24,000 480 0.658 0.701 0.055 1.0
sentence-tone 5 scale 15,000 300 0.637 0.885 0.100 1.26
news-section 4 choice 18,000 360 0.931 0.160 0.027 0.891
tweet-emotion 4 choice 6,514 240 0.879 0.345 0.058 1.059
entailment-short 3 scale 18,000 360 0.906 0.262 0.024 0.944
entailment 3 scale 24,000 480 0.873 0.346 0.027 1.059
tweet-sentiment 3 scale 15,000 300 0.713 0.669 0.049 1.335
spam 2 binary 10,034 240 0.992 0.036 0.005 1.414
product-tone 2 scale 15,000 300 0.960 0.092 0.019 1.059
movie-verdict 2 scale 15,000 300 0.957 0.123 0.031 1.122
paraphrase 2 binary 15,000 300 0.950 0.150 0.029 1.059
short-verdict 2 scale 12,000 240 0.950 0.150 0.016 1.122
answers-question 2 binary 18,000 360 0.942 0.143 0.010 0.749
toxic-comment 2 binary 15,000 300 0.933 0.168 0.024 1.059
follows 2 binary 4,980 240 0.875 0.331 0.043 1.414
same-question 2 binary 18,000 360 0.867 0.298 0.031 1.0
offensive 2 binary 15,000 300 0.860 0.327 0.007 1.122
hateful 2 binary 14,989 300 0.837 0.363 0.051 1.498
same-meaning 2 binary 7,336 240 0.808 0.395 0.051 1.414
grammatical 2 binary 15,000 300 0.803 0.408 0.073 1.587
irony 2 binary 5,724 240 0.787 0.421 0.058 0.841
massive-scenario (unseen) 18 choice — 240 0.742 0.716 0.044 1.189
arxiv-category (unseen) 11 choice — 180 0.456 1.495 0.070 1.587
poem-tone (unseen) 4 scale — 96 0.573 1.081 0.128 1.189
claim-veracity (unseen) 4 choice — 180 0.467 1.222 0.097 1.26
medical-pair (unseen) 2 binary — 180 0.772 0.483 0.074 1.587
subjective (unseen) 2 choice — 180 0.683 0.608 0.035 0.445

Overall: accuracy 0.834, log loss 0.432, Brier 0.228, calibration error 0.009 (95% interval [0.007, 0.016]).

The family

checkpoint encoder download
lenabarretta/sharada-base ModernBERT-base, 150M 300 MB the default, and the one to fine-tune
lenabarretta/sharada-large — you are here ModernBERT-large, 397M 794 MB a few points better, about a third slower
lenabarretta/sharada-multilingual-base mmBERT-base, 308M 616 MB the same model over many more languages
lenabarretta/sharada-multilingual-small mmBERT-small, 141M 282 MB narrower body, for throughput rather than for one fast answer

Everything that makes a model what it is lives in the checkpoint: which encoder it wants, how long a text it reads, the temperature fitted for each task. The library reads all of that out of config.json, so a bigger model or a multilingual one is another repository rather than another version of sharada — and from_pretrained loads any of them with the same line.

That is also why a multilingual model cannot be a flag on this one. This encoder is English down to its tokenizer; reading another language means other weights, not another setting.

The probabilities expire

A temperature is fitted on a distribution, not on a model, so it goes stale when the traffic moves. The run's passport.json records what it was fitted on and what should make you fit it again:

from sharada import check_passport
import json

check_passport(json.load(open("passport.json")), recent_examples, model)
# [] when it finds nothing wrong

Fit it again on your own labelled examples before trusting the numbers on your own traffic — model.calibrate(examples) does it in one call, and model.fit(examples) does it for you.

What it will not do

It does not generate: options or nothing. It reads 256 tokens of text, so longer documents need chunking. It is English: the encoder is English-only and so were the label sets it was trained on — the multilingual checkpoints in the table above cover more. And there is no medicine, law or code in its training mix: the two medical label sets are measured only, and verifying public-health claims lands at the majority-class baseline, which is to say it does not work. Those domains need fine-tuning on your own labelled examples. Finally it is small — where an answer needs a fact that is not in the text in front of it, a frontier model wins.

Apache 2.0.

Downloads last month
104
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lenabarretta/sharada-large

Finetuned
(395)
this model