autotrust/jev-9b-decision-demo

Space

JEV-9B — typed decisions in one forward pass

31

6 commits

updated Sep 26, 2026

See the code

README

JEV-9B — typed decisions in one forward pass

An interactive demo for autotrust/JEV-9B — AutoTrust's first integrated System 1 + System 2 open model, built with the Blocks-of-Experts recipe on a frozen Qwen3.5-9B backbone (Apache-2.0).

System 1 answers typed questions in a single prefill pass — nothing is generated — and returns a calibrated probability distribution:

kindquestionreturns
noul"Is this statement true?"[P(false), P(true)]
choice"Which of these 2–16 options?"one probability per option
score"Where on the ordered 0–5 scale?"distribution over the six levels + expected score

System 2 is the untouched base model: the LoRA adapter is switched off and Qwen3.5-9B generates step-by-step reasoning (thinking mode). The Escalate tab implements the model card's confidence-gated pattern — System 1 answers in one pass, and when its top probability is below your threshold the same weights hand over to System 2.

What to expect

Per the authors' published measurements: mean KL ≈ 0.019 to the closed TypeSafe Jev 1.13 on 25,376 held-out decisions (indistinguishable at that resolution), ECE 0.0007, noul AUROC 0.994, and HumanEval 70.7% on the System 2 path — all 164 completions byte-identical to the base model. System 1 mirrors the teacher including its mistakes (multi-hop reasoning, arithmetic, dates, counting), reads English best, uses the first 1,024 tokens of state, and handles at most 16 options — for longer lists the authors recommend autotrust/JEV-27B. Not for high-stakes decisions; use confidence gating.

Implementation notes

  • Inference follows the model authors' documented transformers + peft path 1:1: backbone with the LoRA adapter on, the 24-slot fp32 decision head applied to the last token's post-norm hidden state, divided by the per-kind calibration temperature (noul 1.002 · choice 0.984 · score 1.012, from calibration.json). System 2 runs with the adapter disabled inside peft.disable_adapter(). The adapter is deliberately left unmerged so both systems live in one process.
  • Weights (18 GB bf16) load once at startup; ZeroGPU streams them into VRAM on each request, so the first call after an idle period is slower than the ~90 ms median the authors measure on a dedicated B200.
  • The demo runs on ZeroGPU (zero-a10g); each decision costs one prefill pass, and System 2 generations are capped by the token-budget slider.

Example data

The gr.Examples rows come from the model authors' own published material (Apache-2.0): the quickstart cases in the model card and the Hacker News front-page / community cases in reports/realworld_9b.json, with expected answers published alongside them.

Credits

Model: AutoTrust AI — Apache-2.0, distilled from the published output distributions of TypeSafe Jev 1.13 (an independent student that shares no weights, code or affiliation with TypeSafe AI). System 2 path: Qwen/Qwen3.5-9B.

gradio
mcp-server

autotrust/jev-9b-decision-demo

Space

JEV-9B — typed decisions in one forward pass

31

6 commits

updated Sep 26, 2026

See the code

README

JEV-9B — typed decisions in one forward pass

An interactive demo for autotrust/JEV-9B — AutoTrust's first integrated System 1 + System 2 open model, built with the Blocks-of-Experts recipe on a frozen Qwen3.5-9B backbone (Apache-2.0).

System 1 answers typed questions in a single prefill pass — nothing is generated — and returns a calibrated probability distribution:

kindquestionreturns
noul"Is this statement true?"[P(false), P(true)]
choice"Which of these 2–16 options?"one probability per option
score"Where on the ordered 0–5 scale?"distribution over the six levels + expected score

System 2 is the untouched base model: the LoRA adapter is switched off and Qwen3.5-9B generates step-by-step reasoning (thinking mode). The Escalate tab implements the model card's confidence-gated pattern — System 1 answers in one pass, and when its top probability is below your threshold the same weights hand over to System 2.

What to expect

Per the authors' published measurements: mean KL ≈ 0.019 to the closed TypeSafe Jev 1.13 on 25,376 held-out decisions (indistinguishable at that resolution), ECE 0.0007, noul AUROC 0.994, and HumanEval 70.7% on the System 2 path — all 164 completions byte-identical to the base model. System 1 mirrors the teacher including its mistakes (multi-hop reasoning, arithmetic, dates, counting), reads English best, uses the first 1,024 tokens of state, and handles at most 16 options — for longer lists the authors recommend autotrust/JEV-27B. Not for high-stakes decisions; use confidence gating.

Implementation notes

  • Inference follows the model authors' documented transformers + peft path 1:1: backbone with the LoRA adapter on, the 24-slot fp32 decision head applied to the last token's post-norm hidden state, divided by the per-kind calibration temperature (noul 1.002 · choice 0.984 · score 1.012, from calibration.json). System 2 runs with the adapter disabled inside peft.disable_adapter(). The adapter is deliberately left unmerged so both systems live in one process.
  • Weights (18 GB bf16) load once at startup; ZeroGPU streams them into VRAM on each request, so the first call after an idle period is slower than the ~90 ms median the authors measure on a dedicated B200.
  • The demo runs on ZeroGPU (zero-a10g); each decision costs one prefill pass, and System 2 generations are capped by the token-budget slider.

Example data

The gr.Examples rows come from the model authors' own published material (Apache-2.0): the quickstart cases in the model card and the Hacker News front-page / community cases in reports/realworld_9b.json, with expected answers published alongside them.

Credits

Model: AutoTrust AI — Apache-2.0, distilled from the published output distributions of TypeSafe Jev 1.13 (an independent student that shares no weights, code or affiliation with TypeSafe AI). System 2 path: Qwen/Qwen3.5-9B.

gradio
mcp-server