Qwen3.5-0.8B-Base-Decision

๐Ÿค— Model ยท ๐Ÿ’ป Source

A prompt-programmable low-latency decision model built on Qwen3.5-0.8B-Base.

v0.1 is an inference interface over Qwen/Qwen3.5-0.8B-Base. It does not ship new weights.

context + question + candidates
        โ†“
   Qwen logits
        โ†“
 candidate softmax
        โ†“
 selected + confidence

Code and serving: github.com/dsif2012/Qwen3.5-0.8B-Base-Decision

Read this first

The base model is pre-trained, not instruction-tuned. Out of the box its decisions are weak and lean toward the first candidate label. v0.1 is the serving path and the starting point for LoRA or distillation, not a finished classifier.

Supported

  • BOOLEAN โ†’ true / false
  • CHOICE โ†’ one label from a dynamic candidate set
  • Full candidate scores, confidence, margin
  • Shared context prefilled once, each question scored on its own copy of that state

Footprint

One RTX 4090, WSL2, Transformers reference kernels, bf16:

  • 1.5 GB VRAM with weights loaded
  • 27k-token context with 16 questions: 2.4 s, 3.7 GB peak VRAM

Example

{
  "context": "...",
  "questions": [
    {
      "id": "can_attack",
      "type": "BOOLEAN",
      "question": "Can the actor attack now?"
    },
    {
      "id": "next_action",
      "type": "CHOICE",
      "question": "What should the actor do next?",
      "choices": ["ATTACK", "CHASE", "WAIT", "RETREAT"]
    }
  ]
}
{
  "can_attack": {"value": true, "confidence": 0.96, "margin": 0.92},
  "next_action": {
    "value": "CHASE",
    "confidence": 0.81,
    "margin": 0.69,
    "ranking": [["CHASE", 0.81], ["ATTACK", 0.12], ["WAIT", 0.05], ["RETREAT", 0.02]]
  }
}

The numbers in this example show the response shape. They are not output from the untuned base model.

Use the base weights

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "Qwen/Qwen3.5-0.8B-Base"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

Decision scoring lives in the GitHub package (pcd). Load Qwen3.5, then POST /decision via pcd-cuda.

License

Apache-2.0. Qwen3.5-0.8B-Base weights remain under the original Qwen Apache-2.0 license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for dsif2012/Qwen3.5-0.8B-Base-Decision

Finetuned
(134)
this model