Inside the JEV Ecosystem: 13 Answer Verifiers on One Test Set

Community Article
Published September 20, 2026

Which system actually verifies LLM answers best — and does the winner help your agent?

Every answer-verification vendor publishes a benchmark, and every one of them wins it. We ran thirteen systems on a single test set of 2,018 items with identical labels, published the scoring code, and put the results side by side.

Typed Decision Leaderboard → https://huggingface.co/spaces/mayafree/typed-decision-leaderboard

The short version: only three systems clear 0.70, the strongest two cannot be separated statistically, and a baseline that looks at nothing but answer length and formatting scores 0.7036 — above eight of the thirteen.


1. The category

A typed decision model returns a structured verdict — a boolean, a choice, or an ordinal score — instead of free text. It runs one forward pass and emits zero generated tokens, which is what makes it one to three orders of magnitude cheaper and faster than asking a frontier model the same question.

TypeSafe AI's Jev established the category under the name System One models, and a small ecosystem has grown around it: open reproductions, architecture-compatible alternatives such as Convai's Laya, and adjacent tools such as Patronus Lynx and Vectara HHEM.

What did not exist was a single test set all of them had been run on.

2. What we measured

Task Given a question and an answer written by some model, score whether the answer is correct
Items 2,018 · 508 incorrect · 5 domains · answers produced by 4 different models
Metric AUC — how well wrong answers sort to the bottom. Threshold-free. 0.5 = coin flip
Axis Factual verification, no grounding document supplied
Aggregation Per domain first, then size-weighted
Uncertainty Paired bootstrap, 3,000 resamples; an interval containing zero yields no rank

Two design choices are worth stating explicitly, because they change what the numbers mean.

Per-domain aggregation. Pooling all 2,018 items into one AUC gives systematically higher numbers, because score scales differ between domains and pooling rewards that difference rather than discrimination. On our own system the pooled figure reads 0.80 where the per-domain mean reads 0.73. We report the lower, reproducible one.

Published baselines. Two reference rows sit in the table: a logistic model over ten surface features of the answer (length, digit count, formatting), and the answering model's own stated confidence. They are there so that every other number has something to be measured against.

3. Results

# System Vendor API Local AUC
1 ZTC (397B) VIDRAFT yes yes 0.7364
2 JEV TypeSafe AI yes no 0.7350
3 ZTC (27B) VIDRAFT yes yes 0.7282
— Length & formatting baseline reference — — 0.7036
4 open-jev 4B pngwn no yes 0.6844
5 Patronus Lynx 8B Patronus AI yes yes 0.5179
6 Laya-Typed-Decisions 421M Convai no yes 0.5144
— Model's own stated confidence reference — — 0.5000
7 Laya-Multilingual 322M Convai no yes 0.4796

Reference — asking an LLM directly (generates tokens, so outside the axis, unranked): GPT-5.2 0.7148 · Qwen3-Next-80B 0.6366 · GPT-4o-mini 0.5878 · Gemini 2.5 Flash-Lite 0.5822.

Different axis: Vectara HHEM-2.1 0.4852. It decides whether an answer follows from a supplied document; this test set supplies none.

The top of the table is a tie

ZTC (397B) leads JEV by 0.0014. The 95% interval on that gap runs from −0.019 to +0.032. Under the board's own rule — an interval containing zero produces no rank — first and second place are not distinguishable, and the page says so above the table rather than in a footnote.

4. Four findings that survive the error bars

4.1 The surface baseline is the real bar

Answer length and formatting alone reach 0.7036. Eight of the thirteen measured systems score below it. A verifier under that line is not detecting correctness; it is detecting shape. Any leaderboard in this category that omits the baseline is flattering its entrants.

4.2 Bigger is not better on this axis

Within the same family, the 397B model beats the 27B overall — but on scientific reasoning the order reverses hard:

scientific reasoning   27B 0.7410   ·   397B 0.6287

A model fourteen times larger scores 0.11 lower. This is consistent with a separate observation from a 180B-class model scoring 0.7146, below a 4B at 0.7284. Verification quality tracks representation geometry, not parameter count.

4.3 "Just ask a frontier model" is expensive and mid-table

GPT-5.2 reaches 0.7148 — below both ZTC configurations and below JEV — while generating tokens and costing roughly $0.55 per 1,000 calls against JEV's $0.024. Smaller judges fall further: GPT-4o-mini 0.5878, Gemini 2.5 Flash-Lite 0.5822.

4.4 An open alternative exists, but the gap is real

Laya is architecturally the closest open analogue to Jev — non-autoregressive, zero generated tokens, Apache-2.0, ~0.015 s per call, which is roughly 140× faster than the hosted API we measured. Its speed claim holds. Its accuracy on this set does not: 0.4796 / 0.5144, against a published claim of beating Jev on the vendor's own English typed-decision benchmark. Both numbers can be true — they are different tests — which is precisely the argument for a shared one.

5. AUC is not the number you deploy on

The ranking answers "which system separates right from wrong answers best." It does not answer "which system helps my agent." We measured the second question directly.

Setup. For each item, score the answer with a verifier; send the lowest-scoring k% to a stronger model to be re-answered; keep everything else. All arms draw from one shared pool of re-answers, so no arm gets luckier retries. Ties are broken randomly and averaged over 200 seeds.

Result at a 20% retry budget, against a 74.83% no-gate baseline:

Gate Final accuracy vs. no gate
ZTC 76.16% +1.34 pp
JEV 74.76% −0.07 pp
Random 74.58% −0.25 pp

The two systems differ by 0.0014 AUC and by 1.4 percentage points of end-to-end agent accuracy.

The mechanism explains it. Re-answering is double-edged:

wrong answers sent back  →  38% get fixed
right answers sent back  →  30% get broken

So a gate's value is set by precision, not recall. At a 20% budget JEV routed 403 items, of which 216 were already correct; ZTC routed 403, of which 195 were already correct. JEV repaired about as much as it damaged. That asymmetry is invisible in an AUC column.

Scope: one escalation target, one item set. We report it as a mechanism, not a universal constant.

6. What we publish, and what we cannot

Published in full: every score, every label, and the grading code.

Not redistributed: the source items. They are drawn from corpora whose licences forbid redistribution or modification, or which are access-gated, and from commercial model outputs whose terms do not clearly permit republication. The grading code is published so the identical protocol can be run against any private set.

Listed without a score: four published reproductions that would not execute from their released artefacts — a missing classifier head, unsupported architectures, an incomplete tokenizer. Each appears on the board with the failure and a link, because a category's real state includes what does not run.

Marked pending: one entry whose re-measurement disagreed with its earlier value by a margin our reproduction harness cannot yet account for. It keeps its prior number, flagged, until the harness matches the upstream implementation. Lowering a competitor's score on the strength of our own untrusted code is not a correction.

7. Reproducing this

The board reports, for every entry: API availability, local weight availability, whether it can be re-fitted on customer data, generated-token count, cost per 1,000 calls, and per-domain AUC.

To add a system, open a discussion on the Space with a link and a runnable scoring snippet. Entries that cannot be executed from published artefacts are listed under did not run rather than omitted.

Leaderboard: https://huggingface.co/spaces/mayafree/typed-decision-leaderboard


Systems measured

Darwin-397B-ZTC · ZTC-Judge-27B · Jev · open-jev 4B · Patronus Lynx 8B · Laya · Vectara HHEM-2.1 · GPT-5.2 · GPT-4o-mini · Gemini 2.5 Flash-Lite · Qwen3-Next-80B

Community

Sign up or log in to comment