bmo-intent 1.0

4B-level intent selection in 0.6B parameters, for a local Russian voice assistant.

Qwen3-Embedding-0.6B with LoRA, distilled from Qwen3-Embedding-4B on 18k teacher-labelled utterances. It picks the intent whose description is closest to the utterance, so new intents need no retraining.

Language: Russian (ru-RU). Utterances, the instruction and intent descriptions are in Russian; English benchmarks below show how it transfers.

bmo-intent 1.0 on MASSIVE ru compared with Qwen3 embedders and rerankers

86.7 accuracy on MASSIVE ru, 60 intents — equal to the 4B teacher
+27.2 points over the base Qwen3-Embedding-0.6B
96.8% accuracy when confidence > 90% (63% of utterances)
1.2 GB fits next to Whisper small on an 8 GB MacBook M1

Comparison

MASSIVE ru test, 1,000 utterances. "Unseen 15" = choosing among the 15 intents never seen in training. BTZSC = mean of AG News, Emotion, Banking77 (English). ECE = calibration error after temperature scaling, lower is better.

Model Size MASSIVE ru Unseen 15 BTZSC typed-decisions ECE ↓
bmo-intent 1.0 0.6B 86.7 95.3 65.3 38.2 3.3
Qwen3-Embedding-0.6B (base) 0.6B 59.5 88.4 65.3 40.7 4.0
multilingual-e5-small, distilled the same way 0.12B 82.4 90.2 53.3 34.4 2.8
rubert-tiny2, distilled the same way 0.03B 81.3 86.0 35.4 36.1 4.5
Qwen3-Embedding-4B + heads (teacher) 4B 86.7 95.7 65.0 — 3.8
Qwen3-Reranker-8B 8B 78.9 96.3 67.7 51.5 4.9
Jev 1.13 (TypeSafe, published numbers) — — — 75.3 72.7 —

bmo-intent numbers are the mean of three seeds (86.5 / 86.4 / 87.1). This repository holds the seed-2 adapter: 87.1, reproduced from these files with the code below.

bmo-intent 1.0 on English benchmarks compared with published Jev numbers

Usage

A LoRA adapter (18 MB) on top of Qwen/Qwen3-Embedding-0.6B. The utterance goes in with the Qwen instruction, each intent as "description (key)"; the vector is the last token, L2-normalised.

import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-Embedding-0.6B")
tok.padding_side = "right"
base = AutoModel.from_pretrained("Qwen/Qwen3-Embedding-0.6B")
model = PeftModel.from_pretrained(base, "bmo-story-lab/bmo-intent-1.0").eval()

# Russian instruction: "Determine the voice-assistant user's intent from the utterance"
TASK = "Определи намерение пользователя голосового ассистента по его реплике"

@torch.no_grad()
def embed(texts):
    enc = tok(texts, padding=True, truncation=True, max_length=512, return_tensors="pt")
    h = model(**enc).last_hidden_state
    v = h[torch.arange(len(texts)), enc["attention_mask"].sum(-1) - 1]  # last token
    return F.normalize(v, dim=-1)

intents = {
    "alarm_set": "поставить будильник, разбудить в определённое время",  # set an alarm, wake me at a time
    "weather_query": "узнать погоду",                                      # check the weather
    "play_music": "включить музыку, песню, исполнителя",                   # play music, a song, an artist
}
q = embed([f"Instruct: {TASK}\nQuery:разбуди меня завтра в семь"])  # "wake me up at seven tomorrow"
o = embed([f"{d} ({k})" for k, d in intents.items()])
probs = (33.2 * q @ o.T).softmax(-1)[0]  # 33.2 = learned logit scale
print(dict(zip(intents, probs.round(decimals=3).tolist())))

Intents are just text: add or rename them without retraining. For calibrated confidence, fit a temperature on a few hundred labelled utterances from your own domain.

Limitations

  • Yes/no questions (typed-decisions): below Qwen3-Reranker-8B and Jev — a bi-encoder barely separates "Yes. This is true: …" from "No. This is false: …".
  • Speech: on Whisper small transcripts (19% WER) accuracy drops by 6.4 points; training on transcripts recovers only 1.1.
  • English: Banking77 (72 labels) is well below Jev; the model was tuned on Russian.
  • Many MASSIVE intents have fewer than 10 test utterances; per-intent numbers are noisy.

Training

Qwen3-Embedding-0.6B, LoRA r=16 in the encoder, fp16 base with fp32 adapters, 2× T4 on Kaggle (~12 min per seed). Loss: KL to Qwen3-Embedding-4B soft labels + cross-entropy to the gold intent. 45 of 60 MASSIVE ru intents in training.

По-русски

Модель выбора намерения для локального голосового ассистента: Qwen3-Embedding-0.6B с LoRA, обученная на разметке Qwen3-Embedding-4B. На MASSIVE ru (60 намерений) — 86,7, как у учителя 4B, при размере 1,2 ГБ. Слабые места — вопросы «да/нет» и распознанная речь (−6,4 пункта на Whisper small). Полный отчёт на русском: bmo-story.tech/ru/lab/bmo-intent.

Links

Apache 2.0 · bmo lab

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bmo-story-lab/bmo-intent-1.0

Adapter
(31)
this model

Dataset used to train bmo-story-lab/bmo-intent-1.0

Evaluation results