ntlm1686/Your-language-model-is-already-a-decision-model

Python

0

0 commits

updated Sep 25, 2026

See the code

See what people are saying

SourceMessageScoreDate

Your Language Model Is Already a Decision Model

1

Sep 25, 2026

README

Your Language Model Is Already a Decision Model

A language model can already be called as a decision model. Give it the state and candidate actions, then choose an action from its next-token option probabilities. No decision-specific training is required.

Results

The Qwen3.5-9B checkpoint received no decision-specific training. Its language-model head selects among the same text candidates sent to hosted Jev 1.13.0. The table uses the same archived items for both models. All Qwen3.5-9B rows use the same optimized scorer, with accuracy and time measured together in the new run. Time means a typical time for one decision, or one complete task for DeepSWE. Qwen runs locally and Jev runs through a hosted API; their timing details are in the reproduction notes.

BenchmarkMetricQwen3.5-9BQwen timeJev 1.13.0Jev time
JevBench publicAccuracy80.5% (186/231)64 ms86.1% (199/231)336 ms
WebPRMBench sampleAccuracy49.2% (59/120)123 ms54.2% (65/120)371 ms
WebPRMBench fullAccuracy59.0% (673/1,141)132 ms66.6% (760/1,141)409 ms
DeepSWEPass rate50.0% (19/38)2,543 ms52.6% (20/38)2,280 ms
PhishNChipsAccuracy73.9% (1,478/2,000)55 ms62.5% (1,251/2,000)336 ms
MetaTool Task1Accuracy83.4% (867/1,040)55 ms77.1% (802/1,040)370 ms
When2CallAccuracy61.8% (371/600)139 ms73.3% (440/600)371 ms
BFCL v4 MultipleAccuracy99.5% (199/200)58 ms99.0% (198/200)365 ms

Accuracy counts decisions matching the benchmark answer. DeepSWE pass rate counts tasks where the selected trajectory passed its recorded verifier. Parentheses show successful items / evaluated items.

Qwen3.5-9B and hosted Jev 1.13.0 Doom replays side by side, showing step, kills, model input, action scores, and latency

Qwen3.5-9B (left) and Jev 1.13.0 (right), both with a 15-second budget on seed 100. The replay follows real elapsed time and finishes 12 : 3. Both panels show TIMEOUT when the 15-second evaluation window ends. If an episode ends earlier, its final score is held until the window closes. Protocol and traces.

Qwen3.5-9B next-token log-probability decision rule

Serve Qwen through a Jev-format API

The Qwen3.5-9B local adapter accepts Jev's POST /v1/systemone JSON interface with state and typed questions. It returns the expected answers for Choice, Noul, and Score, while identifying the actual model as qwen3.5-9b. Its probabilities are uncalibrated. Its JevBench API evaluation scored 186/231. The Qwen3-8B adapter uses the same schema and also matched all 231 direct predictions.

CUDA_VISIBLE_DEVICES=0 QWEN35_USE_FLA=1 models/open_jev_py311/bin/python scripts/serve_qwen35_jev.py \
  --device cuda:0 --port 8796 --max-length 8192 --fallback-margin 0
curl -sS http://127.0.0.1:8796/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.5-9b","state":"Customer asks for a refund.","questions":{"route":{"type":"choice","instructions":"Where should the case go?","criteria":{"returns":"Refunds and exchanges","shipping":"Delivery issues"}}}}'

For Choice requests with up to 26 options, the adapter uses rotated letter scores. Larger sets use separately scored Yes/No logits per candidate; the response metadata identifies the method. See the API protocol and protocol tests.

Full model comparison

Accuracy is shown below; DeepSWE uses pass rate. Both Qwen models received no decision-specific training. Each row compares the same items across models.

Benchmark (items)Qwen3.5-9BQwen3-8BCLMJev 1.13.0Open-Jev-2BOpen-Jev-9B
JevBench public (231)80.5%70.6%41.1%86.1%64.9%77.5%
WebPRMBench sample (93)53.8%57.0%20.4%58.1%46.2%50.5%
WebPRMBench full (842)61.6%62.7%18.1%68.9%50.0%56.1%
DeepSWE (38)50.0%47.4%55.3%52.6%63.2%†60.5%†
PhishNChips (2,000)73.9%66.5%60.6%62.5%50.0%72.7%
MetaTool Task1 (1,040)83.4%81.0%49.8%77.1%66.5%76.2%
When2Call (599)61.8%51.1%33.4%73.3%39.2%75.3%
BFCL v4 Multiple (200)99.5%99.0%84.5%99.0%98.0%98.5%

WebPRM and When2Call use the shared subset that fits Open-Jev's default 4K context. † DeepSWE uses the exploratory Open-Jev 8K runs; neither checkpoint could fit these inputs at 4K. CLM uses its default 2K context except for the separate DeepSWE-specific head. Per-item results and evaluation details include the original counts and context limits.

Qwen3.5-9B beats Qwen3-8B on six of the eight complete benchmark sets. Its WebPRM full-set accuracy is six decisions lower (673 versus 679 on the same 1,141 states). One fixed candidate order was tested on WebPRM, so this small difference does not establish that 8B is a better web decision model.

Closed-loop Doom experiment

Both models receive the same text observations and three actions in ViZDoom Defend the Center, with 15 seconds per episode and seeds 100–104.

ModelKills, seeds 100–104Mean killsTypical decision time
Qwen3.5-9B12, 10, 6, 9, 1410.255 ms
Jev 1.13.03, 3, 3, 3, 22.8336 ms

Time includes the local API request for Qwen and the remote API request for Jev. Each episode is warmed up before its clock starts. The Doom protocol includes per-step inputs, scores, actual episode endings, and earlier runs with the other models.

Decision making without extra model training

We present the state and candidate actions in one prompt, stop at Answer: , and read Qwen3.5-9B's next-token probabilities for the option letters. We rotate the options so each action appears at every letter position, add its log probabilities across rotations, and choose the highest-scoring action. Neither Qwen3.5-9B nor Qwen3-8B receives decision-specific fine-tuning. The procedure generates no explanation tokens and needs no embedding model.

With N candidates, this rule uses N forward passes, each containing the state and all options. That makes it easy to inspect the per-option scores, but cost grows with context length and candidate count. The exact prompt and scoring rule, Qwen3.5-9B runner, and reproduction notes are available.

Possible decision model architectures

Decision model architectures mainly differ in how state and action are connected and interact within the network.

Three designs can score a candidate action given a state. Qwen's next-token method puts the state and every action in one prompt and reads restricted token probabilities, with no new weights. CLM independently embeds states and actions with released contrastive heads, allowing action embeddings to be cached. A proposed RAM-inspired design would feed action tokens into a Transformer. State-encoder features enter its cross-attention layers as keys and values, while the action tokens supply the queries. The Transformer updates the action tokens and outputs a score for each action; it would need training and has not been evaluated here. The architecture note expands the comparison.

Three decision-model scoring architectures

The interface avoids collecting decision labels and training a head. Specialized models can still do better: Jev leads on JevBench and eligible WebPRM states, while Open-Jev-9B leads on the selected When2Call subset. Qwen3.5-9B leads on MetaTool, phishing, and BFCL. Rotating N options requires scoring N prompt orders; the server batches them into one forward. CLM can reuse action embeddings. None of these scores alone establishes the best architecture for a browser agent.

The restricted-letter softmax is a distribution over the supplied options, not a calibrated probability that an action is correct. The released Qwen, CLM, and Open-Jev paths tested in the main table consume text. The state/action formulation could accept images if the encoders, input pipeline, and alignment data support them; the pixel-input Doom run is a separate vision-model experiment, not evidence that these text paths already process images. Our offline tasks and small closed-loop Doom game do not measure browser task success, total agent tokens, or total workflow cost.

Reproduce the saved results

git clone --filter=blob:none https://github.com/ntlm1686/Your-language-model-is-already-a-decision-model.git
cd Your-language-model-is-already-a-decision-model
python scripts/verify_results.py
python scripts/summarize_open_jev.py
python scripts/summarize_qwen35.py
python scripts/summarize_qwen35_optimized.py
python scripts/verify_doom_timed.py

These standard-library scripts recount the committed records without downloading weights. GPU and API rerun procedures are in the benchmark protocol, Qwen Jev-format API notes, hosted Jev procedure, Open-Jev protocol, and Doom protocol. scripts/setup_sources.py fetches pinned upstream revisions; large weights and raw trajectories are excluded from Git.

Layout

scripts/                 Pinned downloads, GPU inference, and offline checks
data/deepswe/            The 38-task list and normalized comparison inputs
results/jevbench/        Per-item Jev/Qwen/CLM outputs for 231 tasks
results/open_jev_2b/     Local Open-Jev-2B per-item outputs and length audit
results/open_jev_9b/     Local Open-Jev-9B per-item outputs and length audit
results/qwen3_5_9b/      Original scores and API checks; optimized/ holds the main results
results/webprm/          Per-item Jev/Qwen/CLM outputs for 120 and 1,141 states
results/deepswe/         Jev/Qwen/CLM trajectory selections and scores
results/phishing/         Per-email Jev/Qwen/CLM outputs and summary
results/metatool/         Tool-use awareness decisions and summary
results/when2call/        Next-response decisions and summary
results/bfcl/             BFCL v4 function-name selections and summary
results/doom/             Five-seed game summaries, decision traces, and GIF replays
docs/                    Benchmark procedures, limits, and analyses
workflows/               Scripts for fetching TypeSafe's public workflow examples

Third-party data and weights remain with their publishers: CLM, JevBench, WebPRMBench, DeepSWE trajectories, PhishNChips, MetaTool, When2Call, and BFCL.

WebShop evaluation in progress

The three-model closed-loop run covers all 500 WebShop test tasks with the full product catalog: Qwen3.5-9B next-token probabilities, hosted Jev, and Open-Jev-9B. Each model follows its own actions under the same search rule and budget. See the protocol and reproduction steps. Results will be added after the run is complete.

ntlm1686/Your-language-model-is-already-a-decision-model

Python

0

0 commits

updated Sep 25, 2026

See the code

See what people are saying

SourceMessageScoreDate

Your Language Model Is Already a Decision Model

1

Sep 25, 2026

README

Your Language Model Is Already a Decision Model

A language model can already be called as a decision model. Give it the state and candidate actions, then choose an action from its next-token option probabilities. No decision-specific training is required.

Results

The Qwen3.5-9B checkpoint received no decision-specific training. Its language-model head selects among the same text candidates sent to hosted Jev 1.13.0. The table uses the same archived items for both models. All Qwen3.5-9B rows use the same optimized scorer, with accuracy and time measured together in the new run. Time means a typical time for one decision, or one complete task for DeepSWE. Qwen runs locally and Jev runs through a hosted API; their timing details are in the reproduction notes.

BenchmarkMetricQwen3.5-9BQwen timeJev 1.13.0Jev time
JevBench publicAccuracy80.5% (186/231)64 ms86.1% (199/231)336 ms
WebPRMBench sampleAccuracy49.2% (59/120)123 ms54.2% (65/120)371 ms
WebPRMBench fullAccuracy59.0% (673/1,141)132 ms66.6% (760/1,141)409 ms
DeepSWEPass rate50.0% (19/38)2,543 ms52.6% (20/38)2,280 ms
PhishNChipsAccuracy73.9% (1,478/2,000)55 ms62.5% (1,251/2,000)336 ms
MetaTool Task1Accuracy83.4% (867/1,040)55 ms77.1% (802/1,040)370 ms
When2CallAccuracy61.8% (371/600)139 ms73.3% (440/600)371 ms
BFCL v4 MultipleAccuracy99.5% (199/200)58 ms99.0% (198/200)365 ms

Accuracy counts decisions matching the benchmark answer. DeepSWE pass rate counts tasks where the selected trajectory passed its recorded verifier. Parentheses show successful items / evaluated items.

Qwen3.5-9B and hosted Jev 1.13.0 Doom replays side by side, showing step, kills, model input, action scores, and latency

Qwen3.5-9B (left) and Jev 1.13.0 (right), both with a 15-second budget on seed 100. The replay follows real elapsed time and finishes 12 : 3. Both panels show TIMEOUT when the 15-second evaluation window ends. If an episode ends earlier, its final score is held until the window closes. Protocol and traces.

Qwen3.5-9B next-token log-probability decision rule

Serve Qwen through a Jev-format API

The Qwen3.5-9B local adapter accepts Jev's POST /v1/systemone JSON interface with state and typed questions. It returns the expected answers for Choice, Noul, and Score, while identifying the actual model as qwen3.5-9b. Its probabilities are uncalibrated. Its JevBench API evaluation scored 186/231. The Qwen3-8B adapter uses the same schema and also matched all 231 direct predictions.

CUDA_VISIBLE_DEVICES=0 QWEN35_USE_FLA=1 models/open_jev_py311/bin/python scripts/serve_qwen35_jev.py \
  --device cuda:0 --port 8796 --max-length 8192 --fallback-margin 0
curl -sS http://127.0.0.1:8796/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.5-9b","state":"Customer asks for a refund.","questions":{"route":{"type":"choice","instructions":"Where should the case go?","criteria":{"returns":"Refunds and exchanges","shipping":"Delivery issues"}}}}'

For Choice requests with up to 26 options, the adapter uses rotated letter scores. Larger sets use separately scored Yes/No logits per candidate; the response metadata identifies the method. See the API protocol and protocol tests.

Full model comparison

Accuracy is shown below; DeepSWE uses pass rate. Both Qwen models received no decision-specific training. Each row compares the same items across models.

Benchmark (items)Qwen3.5-9BQwen3-8BCLMJev 1.13.0Open-Jev-2BOpen-Jev-9B
JevBench public (231)80.5%70.6%41.1%86.1%64.9%77.5%
WebPRMBench sample (93)53.8%57.0%20.4%58.1%46.2%50.5%
WebPRMBench full (842)61.6%62.7%18.1%68.9%50.0%56.1%
DeepSWE (38)50.0%47.4%55.3%52.6%63.2%†60.5%†
PhishNChips (2,000)73.9%66.5%60.6%62.5%50.0%72.7%
MetaTool Task1 (1,040)83.4%81.0%49.8%77.1%66.5%76.2%
When2Call (599)61.8%51.1%33.4%73.3%39.2%75.3%
BFCL v4 Multiple (200)99.5%99.0%84.5%99.0%98.0%98.5%

WebPRM and When2Call use the shared subset that fits Open-Jev's default 4K context. † DeepSWE uses the exploratory Open-Jev 8K runs; neither checkpoint could fit these inputs at 4K. CLM uses its default 2K context except for the separate DeepSWE-specific head. Per-item results and evaluation details include the original counts and context limits.

Qwen3.5-9B beats Qwen3-8B on six of the eight complete benchmark sets. Its WebPRM full-set accuracy is six decisions lower (673 versus 679 on the same 1,141 states). One fixed candidate order was tested on WebPRM, so this small difference does not establish that 8B is a better web decision model.

Closed-loop Doom experiment

Both models receive the same text observations and three actions in ViZDoom Defend the Center, with 15 seconds per episode and seeds 100–104.

ModelKills, seeds 100–104Mean killsTypical decision time
Qwen3.5-9B12, 10, 6, 9, 1410.255 ms
Jev 1.13.03, 3, 3, 3, 22.8336 ms

Time includes the local API request for Qwen and the remote API request for Jev. Each episode is warmed up before its clock starts. The Doom protocol includes per-step inputs, scores, actual episode endings, and earlier runs with the other models.

Decision making without extra model training

We present the state and candidate actions in one prompt, stop at Answer: , and read Qwen3.5-9B's next-token probabilities for the option letters. We rotate the options so each action appears at every letter position, add its log probabilities across rotations, and choose the highest-scoring action. Neither Qwen3.5-9B nor Qwen3-8B receives decision-specific fine-tuning. The procedure generates no explanation tokens and needs no embedding model.

With N candidates, this rule uses N forward passes, each containing the state and all options. That makes it easy to inspect the per-option scores, but cost grows with context length and candidate count. The exact prompt and scoring rule, Qwen3.5-9B runner, and reproduction notes are available.

Possible decision model architectures

Decision model architectures mainly differ in how state and action are connected and interact within the network.

Three designs can score a candidate action given a state. Qwen's next-token method puts the state and every action in one prompt and reads restricted token probabilities, with no new weights. CLM independently embeds states and actions with released contrastive heads, allowing action embeddings to be cached. A proposed RAM-inspired design would feed action tokens into a Transformer. State-encoder features enter its cross-attention layers as keys and values, while the action tokens supply the queries. The Transformer updates the action tokens and outputs a score for each action; it would need training and has not been evaluated here. The architecture note expands the comparison.

Three decision-model scoring architectures

The interface avoids collecting decision labels and training a head. Specialized models can still do better: Jev leads on JevBench and eligible WebPRM states, while Open-Jev-9B leads on the selected When2Call subset. Qwen3.5-9B leads on MetaTool, phishing, and BFCL. Rotating N options requires scoring N prompt orders; the server batches them into one forward. CLM can reuse action embeddings. None of these scores alone establishes the best architecture for a browser agent.

The restricted-letter softmax is a distribution over the supplied options, not a calibrated probability that an action is correct. The released Qwen, CLM, and Open-Jev paths tested in the main table consume text. The state/action formulation could accept images if the encoders, input pipeline, and alignment data support them; the pixel-input Doom run is a separate vision-model experiment, not evidence that these text paths already process images. Our offline tasks and small closed-loop Doom game do not measure browser task success, total agent tokens, or total workflow cost.

Reproduce the saved results

git clone --filter=blob:none https://github.com/ntlm1686/Your-language-model-is-already-a-decision-model.git
cd Your-language-model-is-already-a-decision-model
python scripts/verify_results.py
python scripts/summarize_open_jev.py
python scripts/summarize_qwen35.py
python scripts/summarize_qwen35_optimized.py
python scripts/verify_doom_timed.py

These standard-library scripts recount the committed records without downloading weights. GPU and API rerun procedures are in the benchmark protocol, Qwen Jev-format API notes, hosted Jev procedure, Open-Jev protocol, and Doom protocol. scripts/setup_sources.py fetches pinned upstream revisions; large weights and raw trajectories are excluded from Git.

Layout

scripts/                 Pinned downloads, GPU inference, and offline checks
data/deepswe/            The 38-task list and normalized comparison inputs
results/jevbench/        Per-item Jev/Qwen/CLM outputs for 231 tasks
results/open_jev_2b/     Local Open-Jev-2B per-item outputs and length audit
results/open_jev_9b/     Local Open-Jev-9B per-item outputs and length audit
results/qwen3_5_9b/      Original scores and API checks; optimized/ holds the main results
results/webprm/          Per-item Jev/Qwen/CLM outputs for 120 and 1,141 states
results/deepswe/         Jev/Qwen/CLM trajectory selections and scores
results/phishing/         Per-email Jev/Qwen/CLM outputs and summary
results/metatool/         Tool-use awareness decisions and summary
results/when2call/        Next-response decisions and summary
results/bfcl/             BFCL v4 function-name selections and summary
results/doom/             Five-seed game summaries, decision traces, and GIF replays
docs/                    Benchmark procedures, limits, and analyses
workflows/               Scripts for fetching TypeSafe's public workflow examples

Third-party data and weights remain with their publishers: CLM, JevBench, WebPRMBench, DeepSWE trajectories, PhishNChips, MetaTool, When2Call, and BFCL.

WebShop evaluation in progress

The three-model closed-loop run covers all 500 WebShop test tasks with the full product catalog: Qwen3.5-9B next-token probabilities, hosted Jev, and Open-Jev-9B. Each model follows its own actions under the same search rule and budget. See the protocol and reproduction steps. Results will be added after the run is complete.

Languages

Python

100.0%