Instructions to use bmo-story-lab/bmo-intent-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bmo-story-lab/bmo-intent-1.0 with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3-Embedding-0.6B") model = PeftModel.from_pretrained(base_model, "bmo-story-lab/bmo-intent-1.0") - Notebooks
- Google Colab
- Kaggle
bmo-intent 1.0
4B-level intent selection in 0.6B parameters, for a local Russian voice assistant.
Qwen3-Embedding-0.6B with LoRA, distilled from Qwen3-Embedding-4B on 18k teacher-labelled utterances. It picks the intent whose description is closest to the utterance, so new intents need no retraining.
Language: Russian (ru-RU). Utterances, the instruction and intent descriptions are in Russian; English benchmarks below show how it transfers.
| 86.7 | accuracy on MASSIVE ru, 60 intents — equal to the 4B teacher |
| +27.2 | points over the base Qwen3-Embedding-0.6B |
| 96.8% | accuracy when confidence > 90% (63% of utterances) |
| 1.2 GB | fits next to Whisper small on an 8 GB MacBook M1 |
Comparison
MASSIVE ru test, 1,000 utterances. "Unseen 15" = choosing among the 15 intents never seen in training. BTZSC = mean of AG News, Emotion, Banking77 (English). ECE = calibration error after temperature scaling, lower is better.
| Model | Size | MASSIVE ru | Unseen 15 | BTZSC | typed-decisions | ECE ↓ |
|---|---|---|---|---|---|---|
| bmo-intent 1.0 | 0.6B | 86.7 | 95.3 | 65.3 | 38.2 | 3.3 |
| Qwen3-Embedding-0.6B (base) | 0.6B | 59.5 | 88.4 | 65.3 | 40.7 | 4.0 |
| multilingual-e5-small, distilled the same way | 0.12B | 82.4 | 90.2 | 53.3 | 34.4 | 2.8 |
| rubert-tiny2, distilled the same way | 0.03B | 81.3 | 86.0 | 35.4 | 36.1 | 4.5 |
| Qwen3-Embedding-4B + heads (teacher) | 4B | 86.7 | 95.7 | 65.0 | — | 3.8 |
| Qwen3-Reranker-8B | 8B | 78.9 | 96.3 | 67.7 | 51.5 | 4.9 |
| Jev 1.13 (TypeSafe, published numbers) | — | — | — | 75.3 | 72.7 | — |
bmo-intent numbers are the mean of three seeds (86.5 / 86.4 / 87.1). This repository holds the seed-2 adapter: 87.1, reproduced from these files with the code below.
Usage
A LoRA adapter (18 MB) on top of Qwen/Qwen3-Embedding-0.6B. The utterance goes in with the Qwen instruction, each intent as "description (key)"; the vector is the last token, L2-normalised.
import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-Embedding-0.6B")
tok.padding_side = "right"
base = AutoModel.from_pretrained("Qwen/Qwen3-Embedding-0.6B")
model = PeftModel.from_pretrained(base, "bmo-story-lab/bmo-intent-1.0").eval()
# Russian instruction: "Determine the voice-assistant user's intent from the utterance"
TASK = "Определи намерение пользователя голосового ассистента по его реплике"
@torch.no_grad()
def embed(texts):
enc = tok(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")
h = model(**enc).last_hidden_state
v = h[torch.arange(len(texts)), enc["attention_mask"].sum(-1) - 1] # last token
return F.normalize(v, dim=-1)
intents = {
"alarm_set": "поставить будильник, разбудить в определённое время", # set an alarm, wake me at a time
"weather_query": "узнать погоду", # check the weather
"play_music": "включить музыку, песню, исполнителя", # play music, a song, an artist
}
q = embed([f"Instruct: {TASK}\nQuery:разбуди меня завтра в семь"]) # "wake me up at seven tomorrow"
o = embed([f"{d} ({k})" for k, d in intents.items()])
probs = (33.2 * q @ o.T).softmax(-1)[0] # 33.2 = learned logit scale
print(dict(zip(intents, probs.round(decimals=3).tolist())))
Intents are just text: add or rename them without retraining. For calibrated confidence, fit a temperature on a few hundred labelled utterances from your own domain.
Limitations
- Yes/no questions (typed-decisions): below Qwen3-Reranker-8B and Jev — a bi-encoder barely separates "Yes. This is true: …" from "No. This is false: …".
- Speech: on Whisper small transcripts (19% WER) accuracy drops by 6.4 points; training on transcripts recovers only 1.1.
- English: Banking77 (72 labels) is well below Jev; the model was tuned on Russian.
- Many MASSIVE intents have fewer than 10 test utterances; per-intent numbers are noisy.
Training
Qwen3-Embedding-0.6B, LoRA r=16 in the encoder, fp16 base with fp32 adapters, 2× T4 on Kaggle (~12 min per seed). Loss: KL to Qwen3-Embedding-4B soft labels + cross-entropy to the gold intent. 45 of 60 MASSIVE ru intents in training.
По-русски
Модель выбора намерения для локального голосового ассистента: Qwen3-Embedding-0.6B с LoRA, обученная на разметке Qwen3-Embedding-4B. На MASSIVE ru (60 намерений) — 86,7, как у учителя 4B, при размере 1,2 ГБ. Слабые места — вопросы «да/нет» и распознанная речь (−6,4 пункта на Whisper small). Полный отчёт на русском: bmo-story.tech/ru/lab/bmo-intent.
Links
- Report: bmo-story.tech/lab/bmo-intent · по-русски
- Benchmark data: bmo-story-lab/bmo-lab-results
- Base model: Qwen/Qwen3-Embedding-0.6B · teacher: Qwen/Qwen3-Embedding-4B
- Test set: AmazonScience/massive, ru-RU
- More from us: bmo lab
Apache 2.0 · bmo lab
- Downloads last month
- 30
Model tree for bmo-story-lab/bmo-intent-1.0
Dataset used to train bmo-story-lab/bmo-intent-1.0
Evaluation results
- Accuracy, 60 intents (this checkpoint) on MASSIVE (ru-RU, test)test set self-reported87.100

