Mapika/decider-2b-vision

Model

decider-2b-vision: typed decisions from an image in one forward pass

22

4 commits

1 linked in READMEs

updated Sep 20, 2026

See the code

README

decider-2b-vision: typed decisions from an image in one forward pass

The vision variant of decider-2b (v5 language weights transplanted into the full Qwen3.5-2B vision-language model), fine-tuned so that an image (a photo, a diagram, a game frame) plus a text question with lettered options yields a calibrated probability over the options at a single answer slot. No generation. A 256x240 game frame costs 64 visual tokens. Text-only questions work too, with decider-2b v5's behaviour, including its abstention handling.

Contents: The decider family · Usage · Training · Results · Limitations · Changelog

The decider family

All five repositories share one interface (decider.infer.Decider, POST /v1/systemone in TypeSafe's format) and one readout: the letter logits at an answer slot, softmaxed over the options. Pick by size and input.

modelbaseweightsuse it fornumbers
decider-2b v10Qwen3.5-2B-Base3.5 GB bf16the default: routing, classification, judgments, browser agents; 4 ms per request with CUDA graphs on one GPUregression set 0.805 in-task / 0.755 held-out; live browser 93%; Bespoke suite 0.704
decider-35b-a3b v1Qwen3.5-35B-A3B-Base (3B active)65 GB bf16when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage
decider-35b-a3b-nvfp4the 35B in NVFP419.6 GBthe 35B on Blackwell through vLLM or TensorRT-LLM1.0 to 1.5 points under bf16 on the measured fixtures
decider-0.8bQwen3.5-0.8B-Base1.4 GB bf16the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739)
decider-2b-visionQwen3.5-2B vision-language, v5 text weights4.1 GB bf16decisions from an image plus a question; game framesVisual7W 0.89; Breakout 41 from pixels

Code, data registry, training scripts, the changelog and the per-version history: https://github.com/Mapika/decider.

Usage

import torch
from decider.vision import VisionDecisionModel
from decider.infer import Example, Q
m = VisionDecisionModel("<this repo>", grad_ckpt=False).cuda().eval()
ex = Example("This is a visual question about the image.",
             [Q("What is the person holding?", ["a phone", "a cup", "a book", "nothing"], 0)])
inp = m.prepare([(image, ex)])                  # image: PIL image, numpy array, or PNG bytes; None for text-only
probs = torch.softmax(m.slot_logits(inp), -1)[0, :4]

Training

One epoch (80k examples, 50k with images): game frames from Pong, Breakout, CliffWalking, MiniGrid and Super Mario Bros labelled by scripted policies (rare actions oversampled, plus DAgger frames from an earlier model's own play); multiple-choice image tasks from The Cauldron (A-OKVQA, AI2D, ScienceQA, IconQA, TQA, Raven, Hateful Memes); a replay of the text mixture. Then PPO from pixels on Breakout and Pong (the softmax over action options is the policy). Code: https://github.com/Mapika/decider (decider/vision/).

Results

300 items per task.

taskaccuracyECE
Pong frames (agreement with the RAM-state teacher)0.960.02
Breakout frames0.960.02
Visual7W (held out)0.890.03
A-OKVQA / AI2D / ScienceQA / IconQA / Raven / Hateful Memes0.85 / 0.93 / 0.95 / 0.94 / 0.80 / 0.800.02 to 0.07

Playing from pixels only (no text state), three episodes each: Breakout 41 (the RAM-state teacher scores 22), Pong 3 (teacher 8), CliffWalking -13 (optimal), MiniGrid Empty 0.96 (teacher level); held-out Freeway 0, FrozenLake 0, the harder grid worlds 0 (their scripted teachers also score 0), Mario 1-1 315 px. The previous vision release (v4-based) scored Breakout 16, Pong 8, Freeway 8, BabyAI-GoTo 0.30; this one trades Pong and Freeway for Breakout and for the corrected abstention behaviour.

Limitations

Not a chat model and not a captioner: it answers lettered options at one slot. The text weights inside are decider-2b v5, so the text-only behaviour is that of v5 (its abstention handling, none of the v6 to v10 input shapes, calibration or browser results); a retrain on the current text weights has not been released. Game play from pixels is measured on the Atari, MiniGrid and Mario frames it was trained on plus a few held-out games, three episodes each. English only; calibration is measured on the listed datasets, not on your images.

Changelog

versionwhat changed
current weightsv5 text weights transplanted into the Qwen3.5-2B vision-language model, one epoch on game frames, The Cauldron multiple-choice tasks and a text replay, then PPO from pixels on Breakout and Pong
previous releasev4-based: Breakout 16, Pong 8, Freeway 8, BabyAI-GoTo 0.30

Every decider release is listed in docs/CHANGELOG.md of the GitHub repository.

calibrated
conversational
decision-model
image-text-to-text
one-pass
qwen3_5
safetensors
structured-output
vision

Contributors

Mapika

4 commits

Mapika/decider-2b-vision

Model

decider-2b-vision: typed decisions from an image in one forward pass

22

4 commits

1 linked in READMEs

updated Sep 20, 2026

See the code

README

decider-2b-vision: typed decisions from an image in one forward pass

The vision variant of decider-2b (v5 language weights transplanted into the full Qwen3.5-2B vision-language model), fine-tuned so that an image (a photo, a diagram, a game frame) plus a text question with lettered options yields a calibrated probability over the options at a single answer slot. No generation. A 256x240 game frame costs 64 visual tokens. Text-only questions work too, with decider-2b v5's behaviour, including its abstention handling.

Contents: The decider family · Usage · Training · Results · Limitations · Changelog

The decider family

All five repositories share one interface (decider.infer.Decider, POST /v1/systemone in TypeSafe's format) and one readout: the letter logits at an answer slot, softmaxed over the options. Pick by size and input.

modelbaseweightsuse it fornumbers
decider-2b v10Qwen3.5-2B-Base3.5 GB bf16the default: routing, classification, judgments, browser agents; 4 ms per request with CUDA graphs on one GPUregression set 0.805 in-task / 0.755 held-out; live browser 93%; Bespoke suite 0.704
decider-35b-a3b v1Qwen3.5-35B-A3B-Base (3B active)65 GB bf16when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage
decider-35b-a3b-nvfp4the 35B in NVFP419.6 GBthe 35B on Blackwell through vLLM or TensorRT-LLM1.0 to 1.5 points under bf16 on the measured fixtures
decider-0.8bQwen3.5-0.8B-Base1.4 GB bf16the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739)
decider-2b-visionQwen3.5-2B vision-language, v5 text weights4.1 GB bf16decisions from an image plus a question; game framesVisual7W 0.89; Breakout 41 from pixels

Code, data registry, training scripts, the changelog and the per-version history: https://github.com/Mapika/decider.

Usage

import torch
from decider.vision import VisionDecisionModel
from decider.infer import Example, Q
m = VisionDecisionModel("<this repo>", grad_ckpt=False).cuda().eval()
ex = Example("This is a visual question about the image.",
             [Q("What is the person holding?", ["a phone", "a cup", "a book", "nothing"], 0)])
inp = m.prepare([(image, ex)])                  # image: PIL image, numpy array, or PNG bytes; None for text-only
probs = torch.softmax(m.slot_logits(inp), -1)[0, :4]

Training

One epoch (80k examples, 50k with images): game frames from Pong, Breakout, CliffWalking, MiniGrid and Super Mario Bros labelled by scripted policies (rare actions oversampled, plus DAgger frames from an earlier model's own play); multiple-choice image tasks from The Cauldron (A-OKVQA, AI2D, ScienceQA, IconQA, TQA, Raven, Hateful Memes); a replay of the text mixture. Then PPO from pixels on Breakout and Pong (the softmax over action options is the policy). Code: https://github.com/Mapika/decider (decider/vision/).

Results

300 items per task.

taskaccuracyECE
Pong frames (agreement with the RAM-state teacher)0.960.02
Breakout frames0.960.02
Visual7W (held out)0.890.03
A-OKVQA / AI2D / ScienceQA / IconQA / Raven / Hateful Memes0.85 / 0.93 / 0.95 / 0.94 / 0.80 / 0.800.02 to 0.07

Playing from pixels only (no text state), three episodes each: Breakout 41 (the RAM-state teacher scores 22), Pong 3 (teacher 8), CliffWalking -13 (optimal), MiniGrid Empty 0.96 (teacher level); held-out Freeway 0, FrozenLake 0, the harder grid worlds 0 (their scripted teachers also score 0), Mario 1-1 315 px. The previous vision release (v4-based) scored Breakout 16, Pong 8, Freeway 8, BabyAI-GoTo 0.30; this one trades Pong and Freeway for Breakout and for the corrected abstention behaviour.

Limitations

Not a chat model and not a captioner: it answers lettered options at one slot. The text weights inside are decider-2b v5, so the text-only behaviour is that of v5 (its abstention handling, none of the v6 to v10 input shapes, calibration or browser results); a retrain on the current text weights has not been released. Game play from pixels is measured on the Atari, MiniGrid and Mario frames it was trained on plus a few held-out games, three episodes each. English only; calibration is measured on the listed datasets, not on your images.

Changelog

versionwhat changed
current weightsv5 text weights transplanted into the Qwen3.5-2B vision-language model, one epoch on game frames, The Cauldron multiple-choice tasks and a text replay, then PPO from pixels on Breakout and Pong
previous releasev4-based: Breakout 16, Pong 8, Freeway 8, BabyAI-GoTo 0.30

Every decider release is listed in docs/CHANGELOG.md of the GitHub repository.

calibrated
conversational
decision-model
image-text-to-text
one-pass
qwen3_5
safetensors
structured-output
vision

Contributors

Mapika

4 commits