| Decision Index 0.2.1 (balanced skill) | balanced raw | breadth skill | |
|---|---|---|---|
| autotrust/GEV-26B-Decide, adaptive thinking | 62.48 | 70.66 | 62.00 |
| TypeSafe Jev 1.13 (board) | 57.91 | β | β |
| area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---|---|---|---|---|
| GEV-26B-Decide, adaptive thinking | 0.602 | 0.636 | 0.679 | 0.697 | 0.415 |
How the score was computed (our scoring with the kit's score --edition 0.2.1, not a board entry):
autotrust/jev-decision-index-results
(runs/jev-gemma4-26b-a4b, the weights' previous name). Thinking was tried on five of their benchmarks and is not used
there: it adds little to classification, retrieval and tool selection (see
Outside Knowledge & Reasoning).Details: reports/decision_index_adaptive.json, reports/adaptive_latency_summary.json.
The fastest vision model of the family. Every step below is one System 1 decision (thinking off): a screenshot or a camera image in, a probability for every action out, in a single forward pass on one B200.
Computer use: screenshot β which element to click. A real browser (headless Chromium). Every clickable element gets a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.
95% of 60 random multi-step tasks completed (shop, settings, mail; 3β7 clicks each) in about 85 ms per click: the same success rate as JEV-27B-VL, 3Γ faster. The colour swatches carry no text, so that click is decided from the screenshot alone.
Robot arm: pick and place from a camera image. At every step System 1 looks at the top camera image and answers two questions: is the target left or right of the gripper, and above or below it? The arm moves accordingly and halves its step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation).
61 ms per decision, so a whole pick and place takes 4β8 seconds of model time. It completed 40% of 20 random scenes: close to the target its left/right answers are less precise than JEV-27B-VL's, so more grasps miss. Once grasped, 8 of 9 cubes ended in the tray.
Same scenes and tasks for every model in the family:
| GEV-26B-Decide | JEV-27B-VL | JEV-9B | |
|---|---|---|---|
| computer use: numbered boxes + element text (60 tasks) | 95% | 95% | 95% |
| time per click | β 85 ms | β 260 ms | β 200 ms |
| robot arm: pick and place (20 scenes) | 40% | 75% | 50% |
| time per robot-arm decision | 61 ms | 239 ms | 163 ms |
Demo code: JEV-9B vl/demos/ (set JEV_URL to this
server). Per-episode results: reports/demos/.
GEV-26B-Decide answers typed questions with a calibrated probability for every option; with thinking switched on, it thinks only when it needs to. System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities. One set of weights, one vLLM engine, for text and images.
| what it does | output | |
|---|---|---|
| System 1 | typed decisions: yes/no Β· pick one of 2β256 options Β· rate 0β5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
Adaptive thinking (opt-in: thinking: "auto") | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
| System 2 | the unmodified google/gemma-4-26B-A4B-it, optionally thinking step by step, text and images | text / reasoning |
GEV-26B-Decide was previously published as autotrust/JEV-Gemma4-26B-A4B; the weights are the same.
Two models, two organisations. TypeSafe Jev 1.13 is the hosted, closed model made by TypeSafe AI. autotrust/GEV-26B-Decide is an independent open-weights model built by AutoTrust AI; it is not affiliated with, endorsed by, or a product of TypeSafe AI.
Each puzzle has one correct answer. System 1 answers in one pass; with thinking: "auto", System 2 thinks when System 1
is below 0.8 confidence, and its answer is folded into the probabilities. The videos show puzzles that System 1 got wrong;
the tables below count all puzzles.
Minesweeper. Which hidden cell is certainly safe? System 1 is at chance (21.0 % against 25 %); with thinking, 86.0 %. These thoughts usually reach the 8,192-token budget, and the answer read at that point is still right most of the time.
Connect Four. Which column wins now, or stops the opponent from winning next move? 54.0 % β 99.5 %.
Wordle. Which word still fits all the colour feedback? 52.0 % β 100 %.
Sudoku. Which digit belongs in the highlighted cell? 76.0 % β 99.5 %.
One-move puzzles, 200 generated puzzles per game (threshold 0.8, budget 8,192 thinking tokens; text input, chess with the board image as well):
| game | question (options) | chance | System 1 | adaptive | puzzles that thought |
|---|---|---|---|---|---|
| Minesweeper | which hidden cell is certainly safe (1 safe cell, 3 mines) | 25.0 | 21.0 | 86.0 | 100 % |
| Wordle | which word fits all the feedback (8 words) | 12.5 | 52.0 | 100.0 | 98 % |
| Connect Four | which column wins now or blocks (legal columns) | 14.5 | 54.0 | 99.5 | 95 % |
| Sudoku | which digit goes in the cell (1β9) | 11.1 | 76.0 | 99.5 | 92 % |
| 24 game | which expression equals 24 (6 expressions) | 16.7 | 89.0 | 100.0 | 74 % |
| Maze | first step towards the exit (2β4 directions) | 48.4 | 48.5 | 68.0 | 82 % |
| Chess (Lichess puzzles) | which move mates in one (16 moves, board image + FEN) | 6.3 | 46.5 | 78.0 | β |
Whole games and perception, where thinking helps little or not at all:
| check | System 1 | adaptive |
|---|---|---|
| Snake, one game (image + positions) | 3 food in 18 steps | 11 food in 90 steps (about 21 s per step) |
| Connect Four, 6 full games against a heuristic opponent | 1 win, 5 losses | 1 win, 4 losses, 1 draw |
| 2048, one game | 1,476 points | 1,016 points |
| Flappy Bird, one game | 0 pipes (23 frames) | 0 pipes (38 frames, about 35 s per frame) |
| Quick, Draw!, 320 real sketches, 16 answers (30 / 60 / 100 % of the strokes) | 46.9 / 73.1 / 94.4 % | 40.3 / 72.8 / 95.0 % |
Thinking pays off when the answer can be checked step by step against explicit rules: logic puzzles, tactics,
constraints, arithmetic. It does not help perception (sketches), reflexes (Flappy Bird) or long-horizon play (2048, full
Connect Four games), and every thought costs seconds. Details: reports/thinking_games.json.
think_budget caps the thinking tokens (default: no cap beyond the context window; the evaluations below
used 8,192).The threshold and the mix were chosen on 1,754 questions from six public sets that are not part of the Decision Index (test or validation splits, 300 random questions each; AQuA-RAT has 254):
| set | System 1 | adaptive | thinking on | always think |
|---|---|---|---|---|
| AQuA-RAT (math word problems) | 68.1 | 89.0 | 53.9 % | 90.2 |
| LogiQA (logical reasoning) | 55.3 | 80.3 | 63.7 % | 82.0 |
| StrategyQA (multi-hop yes/no) | 68.0 | 78.7 | 71.0 % | 79.0 |
| MedMCQA (medical) | 66.7 | 71.7 | 52.3 % | 73.7 |
| OpenBookQA (science) | 94.3 | 96.3 | 14.7 % | 95.7 |
| CommonsenseQA | 86.7 | 85.0 | 32.0 % | 83.3 |
| all 1,754 | 73.3 | 83.4 | 47.8 % | 83.8 |
Accuracy in %. Calibration is unchanged: ECE 0.035 for System 1 and 0.035 for adaptive. The adaptive mode reaches 96 % of the always-think gain while thinking on 48 % of the questions. Thinking length: median 2,282 tokens, 90th percentile 7,573; 91.6 % of the thoughts finish within the 8,192-token budget.
All ten benchmarks of the area (33,856 scored requests) ran through the server (POST /v1/decide, thinking: "auto",
threshold 0.8, budget 8,192 thinking tokens) and were scored with the Decision Index kit:
| benchmark | chance | System 1 | adaptive | DI skill, System 1 β adaptive | questions that thought |
|---|---|---|---|---|---|
| GPQA Diamond | 25.0 | 42.9 | 78.6 | 0.238 β 0.714 | 87 % |
| CRUXEval | 37.0 | 67.5 | 90.7 | 0.485 β 0.853 | 42 % |
| CLadder | 50.0 | 71.0 | 86.6 | 0.420 β 0.732 | 51 % |
| GSM8K | 25.0 | 97.6 | 99.1 | 0.969 β 0.989 | 5 % |
| SATA-Bench (case exact) | 1.3 | 34.2 | 35.5 | 0.334 β 0.346 | 21 % |
| MuSR | 37.1 | 67.3 | 67.6 | 0.480 β 0.484 | 51 % |
| ChessBench | 8.2 | 23.7 | 23.6 | 0.169 β 0.168 | 88 % |
| HLE | 16.4 | 8.4 | 17.8 | 0.000 β 0.016 | 88 % |
| MMLU-Pro | 11.1 | 65.0 | 84.6 | 0.607 β 0.827 | 78 % |
| BBH | 31.0 | 75.0 | 92.0 | 0.638 β 0.884 | 60 % |
Accuracy in %; "chance" is the kit's random baseline, and the Decision Index skill rescales accuracy so that chance is 0 (below chance counts as 0). The Knowledge & Reasoning area skill rises from 0.429 to 0.602.
Adaptive thinking with the same settings on five benchmarks of the other areas (run stopped after these):
| benchmark | metric | System 1 | adaptive | requests that thought |
|---|---|---|---|---|
| BFCL | case exact accuracy | 94.4 | 96.0 | 7 % |
| API-Bank | accuracy | 84.3 | 86.0 | 18 % |
| CLINC150+OOS (5,456 of 5,500) | macro-F1 | 93.3 | 94.4 | 10 % |
| ToolRet | nDCG@10 | 66.8 | 66.9 | 84 % |
| BANKING77 | macro-F1 | 88.0 | 85.0 | 19 % |
Small gains on tool calls and intents, none on retrieval, and a loss on BANKING77, whose training split System 1 was trained on: the untrained System 2 often overrules a correct System 1 with an over-confident wrong answer. Use thinking for reasoning questions; keep it off for classification, retrieval and tool routing.
Thinking costs time. Over the area's ten benchmarks, 66 % of the requests thought at least once. Thinking length:
median 8,192 tokens (the budget) on GPQA, HLE and ChessBench, about 5,400 on MMLU-Pro, 3,300 on BBH and 900 on GSM8K.
Estimated single-request latency on one idle B200 (System 1 β 45 ms, thinking β 250 tokens per second): about 0.05 s
without thinking and up to about 33 s with an 8,192-token thought; median 13.4 s over the area (90th percentile 33 s).
With speculative decoding (below) thinking runs at about 438 tokens per second: median 7.7 s, 90th percentile 19 s.
Lower threshold or think_budget to trade accuracy for speed; thinking: "off" keeps every decision in one pass.
MMLU-Pro and BBH (except six questions) ran with speculative decoding, which does not change the output distribution.
Six BBH questions with 18 options hit a vLLM error in the speculative-decoding read-out path and were answered by the same
server without speculative decoding; the current serve_decide.py reads answers through a path that works with
speculative decoding.
The checkpoint contains Gemma-4's vision encoder (no audio encoder), so System 1 and System 2 both accept images. The decision head was trained on text; decisions over images are zero-shot.
| check | result |
|---|---|
| synthetic images: colour (8 options), shape (4), printed number (8), "is there a red object?" (yes/no) | 100 % on each (30 images each) |
| VL-RewardBench, 1,247 pairs, both presentation orders averaged | 78.4 % overall (general 55.8, hallucination 84.9, reasoning 76.0; macro 72.2) |
For reference, autotrust/JEV-27B-VL scores 78.3 % on VL-RewardBench with the same protocol.
The backbone's native context is 262,144 tokens (256K). We tested decisions that hinge on a single sentence placed at a random depth in long real text (concatenated PubMedQA abstracts): a yes/no question and a 16-option question, 10 of each per length, on one B200 with vLLM, one request at a time.
| prompt length | yes/no correct | 16-option correct | mean probability on the right answer | median latency |
|---|---|---|---|---|
| 4K | 10/10 | 10/10 | 0.999 | 0.15 s |
| 32K | 10/10 | 10/10 | 0.999 | 1.6 s |
| 64K | 10/10 | 10/10 | 0.999 | 5.0 s |
| 128K | 10/10 | 10/10 | 1.000 | 17.8 s |
Lengths above 128K have not been tested yet.
choice takes 2β256 options. Up to 16 are read in one pass with the trained labels AβP. More options are read in groups
of at most 16 (in parallel), then a final of 16; every option is read and none is pruned (strategy: "tournament", the
default). Zero-shot intent classification, all options offered at once, 400 test utterances per row:
| test | options | accuracy |
|---|---|---|
| MASSIVE (en) | 59 | 91.2 % |
| BANKING77 | 77 | 81.5 % |
| CLINC150 | 150 | 95.5 % |
| CLINC150 utterances among CLINC150 + BANKING77 + MASSIVE intents | 255 | 89.0 % |
| BANKING77 utterances among the same 255 intents | 255 | 75.2 % |
The 255-option sets merge three catalogues with overlapping intents, so part of the drop comes from near-duplicate labels.
A single pass with labels beyond P (strategy: "single") is about 3Γ faster but less accurate here (CLINC150: 88.5 %
against 95.2 %), so it is not the default. The BANKING77 and CLINC150 training splits are part of the training data (see
below).
hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide
bash GEV-26B-Decide/serve.sh # vLLM on :8000; one GPU with 80 GB or more
serve.sh runs serve_decide.py: the standard vLLM OpenAI server (same flags as vllm serve) with a POST /v1/decide
route. It loads the backbone once: plain requests are System 2, and requests for the LoRA module jev-decision
(adapter_vllm/: backbone LoRA + the decision head as an lm_head LoRA) are System 1. It needs a vLLM build with
Gemma-4 support plus patches/vllm-gemma4-lm-head-lora.patch (LoRA on Gemma-4's tied lm_head, vocabulary 262,144);
tested with a vLLM development build from September 2026.
POST /v1/decidecurl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
"question": "How much does the ball cost?",
"options": ["$0.10", "$0.05", "$1.00", "$0.55"],
"thinking": "auto"}'
| field | value |
|---|---|
kind | noul: yes/no, probabilities for ["false", "true"] Β· score: 0β5 Β· choice: your options |
state | what the decision is about: a string, a JSON object, or a list mixing text and images ["Photo: ", {"image": "https://β¦ or data:β¦"}] |
question | one question about the state |
options | choice only: 2β256 strings |
thinking | "off" (default: System 1 only), "auto" (adaptive), "on" (always think); noul and choice. Switch on "auto" for reasoning questions; keep "off" for classification, retrieval and tool routing |
threshold | System 1 confidence below which "auto" thinks (default 0.8) |
think_budget | maximum thinking tokens; default: no cap beyond the context window |
chat_template_kwargs | passed to the base model's chat template, as in its chat API (Gemma-4 has thinking on/off only, so there is no reasoning_effort setting) |
strategy | more than 16 options: "tournament" (default), "single", "permute" |
return_reasoning / debug | include System 2's reasoning / the System 1 and System 2 distributions |
The response has options, probabilities, choice, choice_index, usage and, when thinking was requested,
thinking: {"used": true, "think_tokens": β¦, "think_seconds": β¦, "finished_within_budget": β¦}. GET /v1/decide/info
lists the defaults.
import requests
def decide(kind, state, question, options=None, thinking="auto"):
body = {"kind": kind, "state": state, "question": question, "thinking": thinking, **({"options": options} if options else {})}
r = requests.post("http://localhost:8000/v1/decide", json=body).json()
return dict(zip(r["options"], r["probabilities"])), r.get("thinking", {}).get("used")
decide("noul", "John was born on 29 February 1996.", "Was John's 7th birthday celebrated on a 29 February?")
requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "autotrust/GEV-26B-Decide",
"messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
"max_tokens": 200, "chat_template_kwargs": {"enable_thinking": False}})
transformers engine below to a mean largest probability difference of 0.015 (300 held-out decisions).Faster thinking with speculative decoding. MTP=1 bash GEV-26B-Decide/serve.sh adds Google's 0.9 GB draft model for
this backbone (--speculative-config '{"model": "google/gemma-4-26B-A4B-it-assistant", "num_speculative_tokens": 4}').
On one B200 it speeds up System 2 about 1.8β1.9Γ: 247 β 438 tokens per second for a single request, 7,072 β 13,505 tokens
per second at 128 concurrent requests (mean acceptance length 3.5β3.7 of 4). System 1 decisions are unchanged, but System 1
throughput at high concurrency drops (257 β 140 decisions per second with 64 clients); use it when you mostly think.
serve_decide.py reads answers with the token restriction that works under speculative decoding, which needs
--max-logprobs 256 (set in serve.sh).
If you call /v1/completions for System 1 yourself, pass top_k: 0 and top_p: 1.0: the model's generation config sets
top_k=64 and top_p=0.95, which vLLM applies as request defaults and which would truncate the returned probabilities.
import json, torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from safetensors.torch import load_file
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
d = snapshot_download("autotrust/GEV-26B-Decide")
tok = AutoTokenizer.from_pretrained(d)
base = Gemma4ForConditionalGeneration.from_pretrained(d, dtype=torch.bfloat16, device_map="cuda")
# System 2: `base` is gemma-4-26B-A4B-it unchanged; use base.generate(...) (text or images).
# System 1: adapter merged in memory + head
m = PeftModel.from_pretrained(base, f"{d}/adapter").merge_and_unload().eval()
backbone = m.model
jc, T = json.load(open(f"{d}/judge_config.json")), json.load(open(f"{d}/calibration.json"))["per_kind"]
head = load_file(f"{d}/head.safetensors"); W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda()
@torch.no_grad()
def decide(kind, state, question, options):
lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
ids = torch.tensor([[tok.bos_token_id] + tok.encode(text, add_special_tokens=False)], device="cuda")
h = backbone(input_ids=ids, use_cache=False).last_hidden_state[0, -1].float()
z = 30.0 * torch.tanh((W @ h + b) / 30.0)
s, _ = jc["slots"]["ranges"][kind]
return dict(zip(options, torch.softmax(z[s:s + len(options)] / T[kind], 0).tolist()))
print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
"Is the customer asking for a refund?", ["false", "true"]))
This path reads up to 16 options per pass; for more, use the server (or read groups of 16 and a final, as above).
bare-v1, prefixed with <bos>: [kind] β¦ [state] β¦ [question] β¦ [options] A) β¦ [decision]:| table | noul | choice | score |
|---|---|---|---|
calibration.json (default; used in the Decision Index run) | 1.003 | 1.017 | 0.999 |
calibration_gold.json (calibrated against ground-truth answers) | 1.214 | 1.098 | 1.000 |
Use calibration_gold.json when you gate automatic actions on confidence.
System 1 was trained on teacher distributions and ground-truth decision data. The ground-truth data includes the public training splits of some datasets whose test splits the Decision Index uses (among them BANKING77 and CLINC150); no test split of any benchmark was used, and suite items were excluded before training. The list has been provided to the Decision Index maintainers. MMMU / MMMU-Pro are not valid evaluations for this model. The adaptive-thinking settings were chosen on data outside the Decision Index suite.
noul and score accept only their canonical options.model-*.safetensors Β· config.json Β· processor_config.json Β· tokenizer* Β· chat_template.jinja Β· generation_config.json
google/gemma-4-26B-A4B-it, unchanged (System 2; text + image input)
adapter/ System 1 LoRA (peft), for the transformers path
head.safetensors 24-slot decision head (fp32): proj.weight [24, 2816], proj.bias [24]
judge_config.json slot layout, verbalizer ids, softcap, read-out
calibration.json per-kind temperatures (default)
calibration_gold.json per-kind temperatures calibrated against ground-truth answers
adapter_vllm/ System 1 for vLLM: backbone LoRA + the head as an lm_head LoRA, plus decision_head.json
serve_decide.py Β· serve.sh
vLLM server with POST /v1/decide (System 1, adaptive thinking) next to the OpenAI endpoints
patches/ vLLM patch: LoRA on Gemma-4's tied lm_head
reports/ adaptive thinking: validation summary, Decision Index recomputation, latency summary, other-area sample, games
reports/demos/ computer use and robot arm: per-episode results
videos/ adaptive-thinking videos (Minesweeper, Connect Four, Wordle, Sudoku); computer use and robot arm
Apache-2.0 for the adapter, head and calibration files; base model under the Gemma 4 terms (https://ai.google.dev/gemma/docs/gemma_4_license). Not affiliated with TypeSafe AI.
| Decision Index 0.2.1 (balanced skill) | balanced raw | breadth skill | |
|---|---|---|---|
| autotrust/GEV-26B-Decide, adaptive thinking | 62.48 | 70.66 | 62.00 |
| TypeSafe Jev 1.13 (board) | 57.91 | β | β |
| area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---|---|---|---|---|
| GEV-26B-Decide, adaptive thinking | 0.602 | 0.636 | 0.679 | 0.697 | 0.415 |
How the score was computed (our scoring with the kit's score --edition 0.2.1, not a board entry):
autotrust/jev-decision-index-results
(runs/jev-gemma4-26b-a4b, the weights' previous name). Thinking was tried on five of their benchmarks and is not used
there: it adds little to classification, retrieval and tool selection (see
Outside Knowledge & Reasoning).Details: reports/decision_index_adaptive.json, reports/adaptive_latency_summary.json.
The fastest vision model of the family. Every step below is one System 1 decision (thinking off): a screenshot or a camera image in, a probability for every action out, in a single forward pass on one B200.
Computer use: screenshot β which element to click. A real browser (headless Chromium). Every clickable element gets a numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.
95% of 60 random multi-step tasks completed (shop, settings, mail; 3β7 clicks each) in about 85 ms per click: the same success rate as JEV-27B-VL, 3Γ faster. The colour swatches carry no text, so that click is decided from the screenshot alone.
Robot arm: pick and place from a camera image. At every step System 1 looks at the top camera image and answers two questions: is the target left or right of the gripper, and above or below it? The arm moves accordingly and halves its step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation).
61 ms per decision, so a whole pick and place takes 4β8 seconds of model time. It completed 40% of 20 random scenes: close to the target its left/right answers are less precise than JEV-27B-VL's, so more grasps miss. Once grasped, 8 of 9 cubes ended in the tray.
Same scenes and tasks for every model in the family:
| GEV-26B-Decide | JEV-27B-VL | JEV-9B | |
|---|---|---|---|
| computer use: numbered boxes + element text (60 tasks) | 95% | 95% | 95% |
| time per click | β 85 ms | β 260 ms | β 200 ms |
| robot arm: pick and place (20 scenes) | 40% | 75% | 50% |
| time per robot-arm decision | 61 ms | 239 ms | 163 ms |
Demo code: JEV-9B vl/demos/ (set JEV_URL to this
server). Per-episode results: reports/demos/.
GEV-26B-Decide answers typed questions with a calibrated probability for every option; with thinking switched on, it thinks only when it needs to. System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities. One set of weights, one vLLM engine, for text and images.
| what it does | output | |
|---|---|---|
| System 1 | typed decisions: yes/no Β· pick one of 2β256 options Β· rate 0β5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
Adaptive thinking (opt-in: thinking: "auto") | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
| System 2 | the unmodified google/gemma-4-26B-A4B-it, optionally thinking step by step, text and images | text / reasoning |
GEV-26B-Decide was previously published as autotrust/JEV-Gemma4-26B-A4B; the weights are the same.
Two models, two organisations. TypeSafe Jev 1.13 is the hosted, closed model made by TypeSafe AI. autotrust/GEV-26B-Decide is an independent open-weights model built by AutoTrust AI; it is not affiliated with, endorsed by, or a product of TypeSafe AI.
Each puzzle has one correct answer. System 1 answers in one pass; with thinking: "auto", System 2 thinks when System 1
is below 0.8 confidence, and its answer is folded into the probabilities. The videos show puzzles that System 1 got wrong;
the tables below count all puzzles.
Minesweeper. Which hidden cell is certainly safe? System 1 is at chance (21.0 % against 25 %); with thinking, 86.0 %. These thoughts usually reach the 8,192-token budget, and the answer read at that point is still right most of the time.
Connect Four. Which column wins now, or stops the opponent from winning next move? 54.0 % β 99.5 %.
Wordle. Which word still fits all the colour feedback? 52.0 % β 100 %.
Sudoku. Which digit belongs in the highlighted cell? 76.0 % β 99.5 %.
One-move puzzles, 200 generated puzzles per game (threshold 0.8, budget 8,192 thinking tokens; text input, chess with the board image as well):
| game | question (options) | chance | System 1 | adaptive | puzzles that thought |
|---|---|---|---|---|---|
| Minesweeper | which hidden cell is certainly safe (1 safe cell, 3 mines) | 25.0 | 21.0 | 86.0 | 100 % |
| Wordle | which word fits all the feedback (8 words) | 12.5 | 52.0 | 100.0 | 98 % |
| Connect Four | which column wins now or blocks (legal columns) | 14.5 | 54.0 | 99.5 | 95 % |
| Sudoku | which digit goes in the cell (1β9) | 11.1 | 76.0 | 99.5 | 92 % |
| 24 game | which expression equals 24 (6 expressions) | 16.7 | 89.0 | 100.0 | 74 % |
| Maze | first step towards the exit (2β4 directions) | 48.4 | 48.5 | 68.0 | 82 % |
| Chess (Lichess puzzles) | which move mates in one (16 moves, board image + FEN) | 6.3 | 46.5 | 78.0 | β |
Whole games and perception, where thinking helps little or not at all:
| check | System 1 | adaptive |
|---|---|---|
| Snake, one game (image + positions) | 3 food in 18 steps | 11 food in 90 steps (about 21 s per step) |
| Connect Four, 6 full games against a heuristic opponent | 1 win, 5 losses | 1 win, 4 losses, 1 draw |
| 2048, one game | 1,476 points | 1,016 points |
| Flappy Bird, one game | 0 pipes (23 frames) | 0 pipes (38 frames, about 35 s per frame) |
| Quick, Draw!, 320 real sketches, 16 answers (30 / 60 / 100 % of the strokes) | 46.9 / 73.1 / 94.4 % | 40.3 / 72.8 / 95.0 % |
Thinking pays off when the answer can be checked step by step against explicit rules: logic puzzles, tactics,
constraints, arithmetic. It does not help perception (sketches), reflexes (Flappy Bird) or long-horizon play (2048, full
Connect Four games), and every thought costs seconds. Details: reports/thinking_games.json.
think_budget caps the thinking tokens (default: no cap beyond the context window; the evaluations below
used 8,192).The threshold and the mix were chosen on 1,754 questions from six public sets that are not part of the Decision Index (test or validation splits, 300 random questions each; AQuA-RAT has 254):
| set | System 1 | adaptive | thinking on | always think |
|---|---|---|---|---|
| AQuA-RAT (math word problems) | 68.1 | 89.0 | 53.9 % | 90.2 |
| LogiQA (logical reasoning) | 55.3 | 80.3 | 63.7 % | 82.0 |
| StrategyQA (multi-hop yes/no) | 68.0 | 78.7 | 71.0 % | 79.0 |
| MedMCQA (medical) | 66.7 | 71.7 | 52.3 % | 73.7 |
| OpenBookQA (science) | 94.3 | 96.3 | 14.7 % | 95.7 |
| CommonsenseQA | 86.7 | 85.0 | 32.0 % | 83.3 |
| all 1,754 | 73.3 | 83.4 | 47.8 % | 83.8 |
Accuracy in %. Calibration is unchanged: ECE 0.035 for System 1 and 0.035 for adaptive. The adaptive mode reaches 96 % of the always-think gain while thinking on 48 % of the questions. Thinking length: median 2,282 tokens, 90th percentile 7,573; 91.6 % of the thoughts finish within the 8,192-token budget.
All ten benchmarks of the area (33,856 scored requests) ran through the server (POST /v1/decide, thinking: "auto",
threshold 0.8, budget 8,192 thinking tokens) and were scored with the Decision Index kit:
| benchmark | chance | System 1 | adaptive | DI skill, System 1 β adaptive | questions that thought |
|---|---|---|---|---|---|
| GPQA Diamond | 25.0 | 42.9 | 78.6 | 0.238 β 0.714 | 87 % |
| CRUXEval | 37.0 | 67.5 | 90.7 | 0.485 β 0.853 | 42 % |
| CLadder | 50.0 | 71.0 | 86.6 | 0.420 β 0.732 | 51 % |
| GSM8K | 25.0 | 97.6 | 99.1 | 0.969 β 0.989 | 5 % |
| SATA-Bench (case exact) | 1.3 | 34.2 | 35.5 | 0.334 β 0.346 | 21 % |
| MuSR | 37.1 | 67.3 | 67.6 | 0.480 β 0.484 | 51 % |
| ChessBench | 8.2 | 23.7 | 23.6 | 0.169 β 0.168 | 88 % |
| HLE | 16.4 | 8.4 | 17.8 | 0.000 β 0.016 | 88 % |
| MMLU-Pro | 11.1 | 65.0 | 84.6 | 0.607 β 0.827 | 78 % |
| BBH | 31.0 | 75.0 | 92.0 | 0.638 β 0.884 | 60 % |
Accuracy in %; "chance" is the kit's random baseline, and the Decision Index skill rescales accuracy so that chance is 0 (below chance counts as 0). The Knowledge & Reasoning area skill rises from 0.429 to 0.602.
Adaptive thinking with the same settings on five benchmarks of the other areas (run stopped after these):
| benchmark | metric | System 1 | adaptive | requests that thought |
|---|---|---|---|---|
| BFCL | case exact accuracy | 94.4 | 96.0 | 7 % |
| API-Bank | accuracy | 84.3 | 86.0 | 18 % |
| CLINC150+OOS (5,456 of 5,500) | macro-F1 | 93.3 | 94.4 | 10 % |
| ToolRet | nDCG@10 | 66.8 | 66.9 | 84 % |
| BANKING77 | macro-F1 | 88.0 | 85.0 | 19 % |
Small gains on tool calls and intents, none on retrieval, and a loss on BANKING77, whose training split System 1 was trained on: the untrained System 2 often overrules a correct System 1 with an over-confident wrong answer. Use thinking for reasoning questions; keep it off for classification, retrieval and tool routing.
Thinking costs time. Over the area's ten benchmarks, 66 % of the requests thought at least once. Thinking length:
median 8,192 tokens (the budget) on GPQA, HLE and ChessBench, about 5,400 on MMLU-Pro, 3,300 on BBH and 900 on GSM8K.
Estimated single-request latency on one idle B200 (System 1 β 45 ms, thinking β 250 tokens per second): about 0.05 s
without thinking and up to about 33 s with an 8,192-token thought; median 13.4 s over the area (90th percentile 33 s).
With speculative decoding (below) thinking runs at about 438 tokens per second: median 7.7 s, 90th percentile 19 s.
Lower threshold or think_budget to trade accuracy for speed; thinking: "off" keeps every decision in one pass.
MMLU-Pro and BBH (except six questions) ran with speculative decoding, which does not change the output distribution.
Six BBH questions with 18 options hit a vLLM error in the speculative-decoding read-out path and were answered by the same
server without speculative decoding; the current serve_decide.py reads answers through a path that works with
speculative decoding.
The checkpoint contains Gemma-4's vision encoder (no audio encoder), so System 1 and System 2 both accept images. The decision head was trained on text; decisions over images are zero-shot.
| check | result |
|---|---|
| synthetic images: colour (8 options), shape (4), printed number (8), "is there a red object?" (yes/no) | 100 % on each (30 images each) |
| VL-RewardBench, 1,247 pairs, both presentation orders averaged | 78.4 % overall (general 55.8, hallucination 84.9, reasoning 76.0; macro 72.2) |
For reference, autotrust/JEV-27B-VL scores 78.3 % on VL-RewardBench with the same protocol.
The backbone's native context is 262,144 tokens (256K). We tested decisions that hinge on a single sentence placed at a random depth in long real text (concatenated PubMedQA abstracts): a yes/no question and a 16-option question, 10 of each per length, on one B200 with vLLM, one request at a time.
| prompt length | yes/no correct | 16-option correct | mean probability on the right answer | median latency |
|---|---|---|---|---|
| 4K | 10/10 | 10/10 | 0.999 | 0.15 s |
| 32K | 10/10 | 10/10 | 0.999 | 1.6 s |
| 64K | 10/10 | 10/10 | 0.999 | 5.0 s |
| 128K | 10/10 | 10/10 | 1.000 | 17.8 s |
Lengths above 128K have not been tested yet.
choice takes 2β256 options. Up to 16 are read in one pass with the trained labels AβP. More options are read in groups
of at most 16 (in parallel), then a final of 16; every option is read and none is pruned (strategy: "tournament", the
default). Zero-shot intent classification, all options offered at once, 400 test utterances per row:
| test | options | accuracy |
|---|---|---|
| MASSIVE (en) | 59 | 91.2 % |
| BANKING77 | 77 | 81.5 % |
| CLINC150 | 150 | 95.5 % |
| CLINC150 utterances among CLINC150 + BANKING77 + MASSIVE intents | 255 | 89.0 % |
| BANKING77 utterances among the same 255 intents | 255 | 75.2 % |
The 255-option sets merge three catalogues with overlapping intents, so part of the drop comes from near-duplicate labels.
A single pass with labels beyond P (strategy: "single") is about 3Γ faster but less accurate here (CLINC150: 88.5 %
against 95.2 %), so it is not the default. The BANKING77 and CLINC150 training splits are part of the training data (see
below).
hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide
bash GEV-26B-Decide/serve.sh # vLLM on :8000; one GPU with 80 GB or more
serve.sh runs serve_decide.py: the standard vLLM OpenAI server (same flags as vllm serve) with a POST /v1/decide
route. It loads the backbone once: plain requests are System 2, and requests for the LoRA module jev-decision
(adapter_vllm/: backbone LoRA + the decision head as an lm_head LoRA) are System 1. It needs a vLLM build with
Gemma-4 support plus patches/vllm-gemma4-lm-head-lora.patch (LoRA on Gemma-4's tied lm_head, vocabulary 262,144);
tested with a vLLM development build from September 2026.
POST /v1/decidecurl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball.",
"question": "How much does the ball cost?",
"options": ["$0.10", "$0.05", "$1.00", "$0.55"],
"thinking": "auto"}'
| field | value |
|---|---|
kind | noul: yes/no, probabilities for ["false", "true"] Β· score: 0β5 Β· choice: your options |
state | what the decision is about: a string, a JSON object, or a list mixing text and images ["Photo: ", {"image": "https://β¦ or data:β¦"}] |
question | one question about the state |
options | choice only: 2β256 strings |
thinking | "off" (default: System 1 only), "auto" (adaptive), "on" (always think); noul and choice. Switch on "auto" for reasoning questions; keep "off" for classification, retrieval and tool routing |
threshold | System 1 confidence below which "auto" thinks (default 0.8) |
think_budget | maximum thinking tokens; default: no cap beyond the context window |
chat_template_kwargs | passed to the base model's chat template, as in its chat API (Gemma-4 has thinking on/off only, so there is no reasoning_effort setting) |
strategy | more than 16 options: "tournament" (default), "single", "permute" |
return_reasoning / debug | include System 2's reasoning / the System 1 and System 2 distributions |
The response has options, probabilities, choice, choice_index, usage and, when thinking was requested,
thinking: {"used": true, "think_tokens": β¦, "think_seconds": β¦, "finished_within_budget": β¦}. GET /v1/decide/info
lists the defaults.
import requests
def decide(kind, state, question, options=None, thinking="auto"):
body = {"kind": kind, "state": state, "question": question, "thinking": thinking, **({"options": options} if options else {})}
r = requests.post("http://localhost:8000/v1/decide", json=body).json()
return dict(zip(r["options"], r["probabilities"])), r.get("thinking", {}).get("used")
decide("noul", "John was born on 29 February 1996.", "Was John's 7th birthday celebrated on a 29 February?")
requests.post("http://localhost:8000/v1/chat/completions", json={
"model": "autotrust/GEV-26B-Decide",
"messages": [{"role": "user", "content": "In one sentence, what is safety stock?"}],
"max_tokens": 200, "chat_template_kwargs": {"enable_thinking": False}})
transformers engine below to a mean largest probability difference of 0.015 (300 held-out decisions).Faster thinking with speculative decoding. MTP=1 bash GEV-26B-Decide/serve.sh adds Google's 0.9 GB draft model for
this backbone (--speculative-config '{"model": "google/gemma-4-26B-A4B-it-assistant", "num_speculative_tokens": 4}').
On one B200 it speeds up System 2 about 1.8β1.9Γ: 247 β 438 tokens per second for a single request, 7,072 β 13,505 tokens
per second at 128 concurrent requests (mean acceptance length 3.5β3.7 of 4). System 1 decisions are unchanged, but System 1
throughput at high concurrency drops (257 β 140 decisions per second with 64 clients); use it when you mostly think.
serve_decide.py reads answers with the token restriction that works under speculative decoding, which needs
--max-logprobs 256 (set in serve.sh).
If you call /v1/completions for System 1 yourself, pass top_k: 0 and top_p: 1.0: the model's generation config sets
top_k=64 and top_p=0.95, which vLLM applies as request defaults and which would truncate the returned probabilities.
import json, torch
from huggingface_hub import snapshot_download
from peft import PeftModel
from safetensors.torch import load_file
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
d = snapshot_download("autotrust/GEV-26B-Decide")
tok = AutoTokenizer.from_pretrained(d)
base = Gemma4ForConditionalGeneration.from_pretrained(d, dtype=torch.bfloat16, device_map="cuda")
# System 2: `base` is gemma-4-26B-A4B-it unchanged; use base.generate(...) (text or images).
# System 1: adapter merged in memory + head
m = PeftModel.from_pretrained(base, f"{d}/adapter").merge_and_unload().eval()
backbone = m.model
jc, T = json.load(open(f"{d}/judge_config.json")), json.load(open(f"{d}/calibration.json"))["per_kind"]
head = load_file(f"{d}/head.safetensors"); W, b = head["proj.weight"].cuda(), head["proj.bias"].cuda()
@torch.no_grad()
def decide(kind, state, question, options):
lines = options if kind != "choice" else [f"{'ABCDEFGHIJKLMNOP'[i]}) {o}" for i, o in enumerate(options)]
text = f"[kind] {kind}\n[state] {state}\n[question] {question}\n[options]\n" + "\n".join(lines) + "\n[decision]:"
ids = torch.tensor([[tok.bos_token_id] + tok.encode(text, add_special_tokens=False)], device="cuda")
h = backbone(input_ids=ids, use_cache=False).last_hidden_state[0, -1].float()
z = 30.0 * torch.tanh((W @ h + b) / 30.0)
s, _ = jc["slots"]["ranges"][kind]
return dict(zip(options, torch.softmax(z[s:s + len(options)] / T[kind], 0).tolist()))
print(decide("noul", "Customer says the parcel arrived damaged and wants their money back.",
"Is the customer asking for a refund?", ["false", "true"]))
This path reads up to 16 options per pass; for more, use the server (or read groups of 16 and a final, as above).
bare-v1, prefixed with <bos>: [kind] β¦ [state] β¦ [question] β¦ [options] A) β¦ [decision]:| table | noul | choice | score |
|---|---|---|---|
calibration.json (default; used in the Decision Index run) | 1.003 | 1.017 | 0.999 |
calibration_gold.json (calibrated against ground-truth answers) | 1.214 | 1.098 | 1.000 |
Use calibration_gold.json when you gate automatic actions on confidence.
System 1 was trained on teacher distributions and ground-truth decision data. The ground-truth data includes the public training splits of some datasets whose test splits the Decision Index uses (among them BANKING77 and CLINC150); no test split of any benchmark was used, and suite items were excluded before training. The list has been provided to the Decision Index maintainers. MMMU / MMMU-Pro are not valid evaluations for this model. The adaptive-thinking settings were chosen on data outside the Decision Index suite.
noul and score accept only their canonical options.model-*.safetensors Β· config.json Β· processor_config.json Β· tokenizer* Β· chat_template.jinja Β· generation_config.json
google/gemma-4-26B-A4B-it, unchanged (System 2; text + image input)
adapter/ System 1 LoRA (peft), for the transformers path
head.safetensors 24-slot decision head (fp32): proj.weight [24, 2816], proj.bias [24]
judge_config.json slot layout, verbalizer ids, softcap, read-out
calibration.json per-kind temperatures (default)
calibration_gold.json per-kind temperatures calibrated against ground-truth answers
adapter_vllm/ System 1 for vLLM: backbone LoRA + the head as an lm_head LoRA, plus decision_head.json
serve_decide.py Β· serve.sh
vLLM server with POST /v1/decide (System 1, adaptive thinking) next to the OpenAI endpoints
patches/ vLLM patch: LoRA on Gemma-4's tied lm_head
reports/ adaptive thinking: validation summary, Decision Index recomputation, latency summary, other-area sample, games
reports/demos/ computer use and robot arm: per-episode results
videos/ adaptive-thinking videos (Minesweeper, Connect Four, Wordle, Sudoku); computer use and robot arm
Apache-2.0 for the adapter, head and calibration files; base model under the Gemma 4 terms (https://ai.google.dev/gemma/docs/gemma_4_license). Not affiliated with TypeSafe AI.