autotrust/JEV-35B

AutoTrust/GuruSearch recommendation & search live demo: news.guru.so

🔴 Live: AutoTrust/GuruSearch recommendation & search demo → news.guru.so

Try AutoTrust/GuruSearch live (new, 11 October 2026). A recommendation and search demo in which every ranking is a calibrated System 1 decision, made in about half a second:

  • Recommendation: the latest headlines from 6 news feeds, ranked by importance.
  • Search: results from Google News, Bing News and Yahoo News re-ranked by relevance, side by side with the search engines' own order, plus a short answer with citations.

A System 1 decision model on Qwen3.5-35B-A3B (35 B parameters, about 3 B active per token), up to 256 options in one pass

Decision Index 0.3, public suite: 59.75 (our scoring with the kit), against 53.64 for autotrust/JEV-27B-VL on the board. Approximate public vision score 72.08 (our rebuild of the vision benchmarks; JEV-27B-VL 71.67 on the same rebuild, a tie within noise). Median single-request latency 241 ms on one B200 (JEV-27B-VL 271 ms).

NVFP4 version: autotrust/JEV-35B-NVFP4, the same model with its routed experts in NVFP4: 25.6 GB download, 23 GiB of GPU memory (bf16: 72 GB, 66 GiB), so it runs on one 32–48 GB GPU. Decision Index 0.3 public 59.18 (bf16 59.75), vision rebuild 71.23 (bf16 72.08), 96.6 % of the Decision Index answers identical to bf16, same latency, same serve.sh and API.

version download GPU memory (weights) Decision Index 0.3 public vision (rebuild)
autotrust/JEV-35B (bf16, this repo) 72 GB 66.5 GiB 59.75 72.08
autotrust/JEV-35B-NVFP4 25.6 GB 23.3 GiB 59.18 71.23

Two models, two organisations. TypeSafe Jev 1.13 is the hosted, closed model made by TypeSafe AI. autotrust/JEV-35B is an independent open-weights model built by AutoTrust AI; it is not affiliated with, endorsed by, or a product of TypeSafe AI.

Decision Index 0.3 (public suite)

public index
autotrust/JEV-35B, System 1 (our run with the kit) 59.75
autotrust/JEV-27B-VL, System 1 (board, public part) 53.64
area (skill) Knowledge & Reasoning Language Retrieval & Classification Tools & Automation Arts & Taste
JEV-35B 0.426 0.632 0.710 0.765 0.422
JEV-27B-VL 0.414 0.564 0.537 0.735 0.416

JEV-35B: all 140,178 scoreable requests of the 0.3 public suite answered (0 errors), System 1 only (thinking off), every choice read in one pass (up to 256 options), scored with the kit's score --edition 0.3. Our scoring, not a board entry: the board's Full score also counts private tests (80 %), which only its maintainers run. JEV-27B-VL: the board's public numbers for JEV-27B, whose text decisions JEV-27B-VL reproduces (see the JEV-27B-VL card).

JEV-35B is ahead in all five areas and on 24 of 37 benchmarks. Largest gains (skill points): PhishNChips +42.3, HoVer +28.3, Habermas Machine +28.1, VAST +22.8, WinoGrande +18.9, BANKING77 +18.3, iSarcasmEval +17.2, When2Call +14.9, GPQA +9.5. JEV-27B-VL is ahead on POP909-CL (+25.7), GSM8K (+18.3), ANLI (+11.4), BPoMP (+10.8), NLI4CT (+6.7) and BBH (+4.8).

Every benchmark

Skill rescales the benchmark's own metric so that chance is 0 (below chance counts as 0); ★ = gold benchmark (weight 1.2); the higher score is in bold.

Knowledge & Reasoning (area skill 0.426; JEV-27B-VL 0.414)

benchmark JEV-35B JEV-27B-VL
GPQA Diamond ★ 0.360 0.265
GSM8K (0.3 rebuild) 0.359 0.542
ChessBench 0.127 0.091
MuSR 0.361 0.402
SATA-Bench 0.284 0.328
CRUXEval 0.546 0.583
CLadder 0.440 0.410
HLE ★ 0.000 0.000
MMLU-Pro ★ 0.642 0.567
BBH ★ 0.647 0.695
WinoGrande ★ 0.841 0.651

Language Understanding (area skill 0.632; JEV-27B-VL 0.564)

benchmark JEV-35B JEV-27B-VL
ContractNLI 0.732 0.639
ANLI ★ 0.545 0.658
HellaSwag ★ 0.962 0.900
ACOS 0.280 0.169
FinEntity 0.856 0.755
iSarcasmEval 0.426 0.254
VAST 0.550 0.322
NLI4CT 0.627 0.694
RAGTruth 0.666 0.604

Retrieval & Classification (area skill 0.710; JEV-27B-VL 0.537)

benchmark JEV-35B JEV-27B-VL
BANKING77 ★ 0.922 0.739
CLINC150+OOS ★ 0.943 0.845
BRIGHT ★ 0.421 0.417
Amazon ESCI 0.519 0.430
PhishNChips 0.661 0.238
HoVer 0.761 0.478

Tools & Automation (area skill 0.765; JEV-27B-VL 0.735)

benchmark JEV-35B JEV-27B-VL
BFCL ★ 0.950 0.946
ToolRet 0.600 0.604
API-Bank ★ 0.856 0.825
Home appliance simulator 0.500 0.523
When2Call 0.863 0.714

Arts & Human Taste (area skill 0.422; JEV-27B-VL 0.416)

benchmark JEV-35B JEV-27B-VL
BPoMP 0.750 0.859
Humicroedit 0.232 0.223
POP909-CL 0.132 0.390
cfcolor 0.256 0.283
Habermas Machine 0.416 0.135
New Yorker 0.744 0.607

Images (approximate public vision score)

The vision benchmarks of the Decision Index are not published; this is our rebuild of their public datasets (CV-Bench, BLINK, RealWorldQA, CharXiv, InfographicVQA, Mind2Web, CORD + FUNSD, Hateful Memes, R-Bench-M, MMMU-Pro vision; Winoground not included), scored with the board's chance correction and weights. The same rebuild gives JEV-27B-VL 71.67 against its board score of 71.53 on the same benchmarks.

approximate public vision score accuracy ECE
JEV-35B 72.08 78.6 % 0.050
JEV-27B-VL (same rebuild) 71.67 78.0 % 0.044
benchmark JEV-35B skill JEV-27B-VL skill
CV-Bench 75.7 76.9
BLINK 55.6 56.7
RealWorldQA 68.9 72.2
CharXiv 80.7 80.7
InfographicVQA 95.0 95.9
Mind2Web 83.6 83.2
KIE (CORD+FUNSD) 98.7 98.6
Moderation (Hateful Memes) 47.7 38.0
R-Bench-M 28.1 30.3
MMMU-Pro vision 43.5 39.2

Overall the two models are level: the difference (+0.4) is inside the noise (paired bootstrap 95% interval −0.7 to +1.6). Only two per-benchmark differences are statistically significant (paired McNemar test, p < 0.01), both in favour of JEV-35B: moderation (+9.6) and MMMU-Pro (+4.3). The small deficits on RealWorldQA, CV-Bench, BLINK, InfographicVQA and R-Bench-M (−1 to −3) are not significant. The vision tower is Qwen3.5-35B-A3B's, unchanged; image decisions are zero-shot.

Speed

One sequential client, the same 587 rows (387 text rows sampled across the Decision Index, 200 image rows), POST /v1/decide on serve.sh (vLLM, System 1 as a LoRA), one B200:

median (p95) text image all
JEV-35B 215 ms (415) 378 ms (784) 241 ms (682)
JEV-27B-VL (same benchmark) 207 ms (617) 457 ms (910) 271 ms (746)

Computer use, robot arm and games

Same demo code, seeds, scenes and opponents as the JEV-27B-VL card; every step is one System 1 decision (POST /v1/decide, thinking off), one model on one B200.

Computer use: screenshot → which element to click. A real browser (headless Chromium). Every clickable element gets a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.

Robot arm: pick and place from a camera image (MuJoCo). At every step System 1 answers two questions from the top camera: is the target left or right of the gripper, and above or below it? The arm halves its step whenever an answer flips.

JEV-35B JEV-27B-VL
Computer use, numbered boxes + element text (60 tasks: shop, settings, mail) 95% 95%
Computer use, numbered boxes only (60 tasks) 38% 10%
ms per click decision, median (6 browsers in parallel) 397 720
Robot arm, binary-decision servo (20 scenes): pick-and-place success 75% 75%
Robot arm: median distance from the cube centre when grasping 2.5 cm 2.7 cm
Robot arm: placed in the tray, once grasped 15/15 15/15
Robot arm: ms per decision, median 239 239
Robot arm, direct choice among 8 motor actions (10 scenes) 0% 0%
  • Computer use with element text: 95%, the same three failures as JEV-27B-VL (shop seeds 12, 15, 20: the colour swatch carries no text and is skipped).
  • Numbered boxes only (every element read from pixels): 38% against 10%. JEV-35B completes 90% of the settings tasks but only 15% of mail and 10% of shop; most failures still declare the task complete too early (24 of 37).
  • Robot arm: same success rate, different scenes. Each model misses 5 of 20 grasps (both miss scenes 4 and 19), every miss 3 cm or more off the cube centre; once grasped, every cube reaches the tray.
  • Choosing directly among 8 motor commands fails for both models; decompose control into simple visual questions.

Games (seed 0 for both models; board as image + text):

game JEV-35B JEV-27B-VL
2048: score / largest tile 336 / 32 2,080 / 128
Connect Four against a rule-based opponent (6 games) 0 wins, 6 losses 0 wins, 6 losses
Flappy Bird: pipes passed 0 3
Snake: food eaten 11 21
Quick, Draw! top-1 among 16 (320 sketches; 30% / 60% / 100% of strokes) 35.0 / 57.5 / 86.3% 41.6 / 62.2 / 87.8%
Chess mate in one, choice among 16 moves (200 Lichess puzzles) 47.0% 52.0%

For game play, use JEV-27B-VL. Over five seeds JEV-35B's largest 2048 tile averages 64 (random play: 102), it passes no Flappy Bird pipe and eats 3–18 pieces of food in Snake (mean 10). Per-episode results: reports/demos/.

Quick start (vLLM)

hf download autotrust/JEV-35B --local-dir JEV-35B
bash JEV-35B/serve.sh          # vLLM on :8000; one GPU with 80 GB or more

On a smaller GPU, use autotrust/JEV-35B-NVFP4 (one GPU with 32 GB or more; same serve.sh and API):

hf download autotrust/JEV-35B-NVFP4 --local-dir JEV-35B-NVFP4
bash JEV-35B-NVFP4/serve.sh

serve.sh runs serve_decide.py: the standard vLLM OpenAI server (same flags as vllm serve) with a POST /v1/decide route. Plain requests are System 2; requests for the LoRA module jev-decision (adapter_vllm/: backbone LoRA + the decision head as an lm_head LoRA) are System 1. Tested with a vLLM development build from September 2026.

The adapter has no LoRA on the routed experts. serve.sh sets JEV_DECIDE_NO_MOE_LORA=1 and --lora-target-modules so that the routed experts run on vLLM's normal MoE kernels (the decision head needs --max-lora-rank 320; vLLM's MoE-LoRA kernel supports at most rank 128).

System 1: POST /v1/decide

curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
  "kind": "choice",
  "state": "Customer message: my card was charged twice for the same order.",
  "question": "Which team should handle this ticket?",
  "options": ["billing", "shipping", "technical support", "account security"]}'
field value
kind noul: yes/no, probabilities for ["false", "true"] · score: 0–5 · choice: your options
state what the decision is about: a string, a JSON object, or a list mixing text and images
question one question about the state
options choice only: 2–256 strings
thinking "off" (default)

The response has options, probabilities, choice, choice_index and usage. GET /v1/decide/info lists the defaults.

System 2

import requests
requests.post("http://localhost:8000/v1/chat/completions", json={
    "model": "autotrust/JEV-35B",
    "messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
    "max_tokens": 200})

If you call /v1/completions for System 1 yourself, pass top_k: 0 and top_p: 1.0 so that the returned probabilities are not truncated.

Engine and read-out

System 1 reads the hidden state at the last token of a bare-text prompt:

[kind] choice
[state] ...
[question] ...
[options]
A) ...
B) ...
[decision]:

A 264-slot linear head (fp32) gives one logit per slot; the active slots of the question's kind are soft-maxed with a per-kind temperature (calibration.json). In vLLM the head is expressed as an lm_head LoRA (adapter_vllm/, rank 320) and the head bias is added client-side (adapter_vllm/decision_head.json).

Limitations

  • System 1 only on this card: thinking (adaptive System 2) has not been evaluated for this model.
  • Knowledge & Reasoning is the weak area (GPQA 0.360, MMLU-Pro 0.642, HLE below chance); JEV-27B-VL is better on GSM8K, ANLI and BBH.
  • Weaker than JEV-27B-VL at sequential game play (2048, Flappy Bird, Snake).
  • The vision score is our rebuild, not the board's; image decisions are zero-shot.
  • English-centric; not for high-stakes decisions without confidence gating.

Files

model.safetensors-* · config.json · tokenizer* · chat_template.jinja · generation_config.json · *_config.json
                          Qwen/Qwen3.5-35B-A3B, unchanged (System 2)
adapter/                  System 1 LoRA (peft), rank 32
head.safetensors          264-slot decision head (fp32)
judge_config.json         slot layout, verbalizer ids, read-out
calibration.json          per-kind temperatures
adapter_vllm/             System 1 for vLLM: backbone LoRA + the head as an lm_head LoRA (rank 320), plus decision_head.json
serve_decide.py · serve.sh
                          vLLM server with POST /v1/decide next to the OpenAI endpoints
videos/                   computer-use and robot-arm episodes
reports/demos/            per-episode computer-use, robot-arm and game results

NVFP4 weights (routed experts in NVFP4, everything else identical): autotrust/JEV-35B-NVFP4.

License

Apache-2.0. This repository contains the weights of Qwen/Qwen3.5-35B-A3B (Apache-2.0) unchanged, plus the System 1 adapter, decision head and calibration.

Downloads last month
360
Safetensors
Model size
36B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for autotrust/JEV-35B

Adapter
(46)
this model
Quantizations
1 model

Evaluation results

  • Decision Index 0.3 public index on Decision Index 0.3 public suite
    self-reported
    59.750