1st of 18 models in the 3–6B class of the Decision Index live board (44.97), and 2nd of 27 when open submissions are included.
jiwo-4b is a decision model. It reads a state and one or more typed questions. For each question, it returns a probability for every option in one forward pass. It does not generate text.
The request and response format is the format of TypeSafe's Jev API (POST /v1/systemone). jiwo is an independent
project. TypeSafe does not endorse it.
| Base model | Qwen/Qwen3.5-4B |
| Parameters | 4.2B |
| Weights | Full weights: decoder, tokenizer and readout |
| Question types | choice (at most 255 options), score (2 to 10 levels), noul (true or false) |
| Training context length | 8,192 tokens |
| Licence | Apache-2.0 |
Code, server and documentation: github.com/jiwidi/jiwo. The other model of the family is eljiwo/jiwo-0.8b.
| Benchmark | Result |
|---|---|
| Decision Index 0.2.1 | 44.97, 1st of 18 in the 3–6B class on the live board of 2026-09-28, and 2nd of 27 with open submissions¹ |
| Area skill ×100: knowledge / language / retrieval and classification / tools / arts | 26.2 / 49.1 / 51.6 / 67.6 / 28.2 |
| Untrained Qwen3.5-4B, same server and suite, temperatures fitted on the same calibration data | 29.20: +15.8 points for jiwo-4b |
| Best other Qwen3.5-4B fine-tune on the live board | JPT-4B, 43.04: +4.5% for jiwo-4b |
| Best other Qwen3.5-4B fine-tune, including open submissions¹ | ezjev-4b-s2, 51.15: −12.1% for jiwo-4b |
| Median request latency on one H100 | 53 ms |
The score comes from a complete run of the suite (150,317 requests), served by jiwo serve with
JIWO_MAX_LENGTH=65536. The board maintainers did not verify it yet.
¹ Complete runs in open pull requests (as of 2026-10-04) with a readable scores.json and a median latency under 1,000 ms, the limit of the board. One slower Qwen3.5-4B submission also scores higher: Wald-Q4B v1.1 (54.59, median 2,360 ms).
pip install "jiwo @ git+https://github.com/jiwidi/jiwo" # on NVIDIA GPUs: "jiwo[cuda] @ git+..."
from jiwo.model import DecisionModel
model = DecisionModel.from_pretrained("eljiwo/jiwo-4b")
question = {"type": "noul", "instructions": "The customer is angry."}
response = model.decide("Refund please, the parcel arrived crushed.", {"angry": question})
print(response["answers"]["angry"]["noul"]) # the probability that the statement is true
Or serve it over HTTP. The server listens on 127.0.0.1:8765:
JIWO_CHECKPOINT=eljiwo/jiwo-4b jiwo serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"state": "Refund please, the parcel arrived crushed.",
"questions": {
"team": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "Payments and refunds", "shipping": "Damaged or lost parcels", "account": null}},
"angry": {"type": "noul", "instructions": "The customer is angry."},
"urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Not urgent", "Soon", "Now"]}}}'
The jiwo README describes the response, the errors and all settings.
A request has a state (a string, a JSON object or a JSON array) and at most 64 named questions.
| Type | Meaning | criteria |
|---|---|---|
choice | Pick one of N named options | An object that maps option keys to descriptions or null |
score | Rate the state on 2 to 10 ordered levels | A list of level descriptions, lowest level first |
noul | Decide whether a statement is true | Optional. Descriptions for the keys true and false |
The model was fine-tuned on the train splits of public datasets and on synthetic data. The training data is mostly English.
JIWO_MAX_LENGTH, but the model did not see longer prompts in training.choice question can have at most 255 options.Apache-2.0, the licence of the base model. See LICENSE.
1st of 18 models in the 3–6B class of the Decision Index live board (44.97), and 2nd of 27 when open submissions are included.
jiwo-4b is a decision model. It reads a state and one or more typed questions. For each question, it returns a probability for every option in one forward pass. It does not generate text.
The request and response format is the format of TypeSafe's Jev API (POST /v1/systemone). jiwo is an independent
project. TypeSafe does not endorse it.
| Base model | Qwen/Qwen3.5-4B |
| Parameters | 4.2B |
| Weights | Full weights: decoder, tokenizer and readout |
| Question types | choice (at most 255 options), score (2 to 10 levels), noul (true or false) |
| Training context length | 8,192 tokens |
| Licence | Apache-2.0 |
Code, server and documentation: github.com/jiwidi/jiwo. The other model of the family is eljiwo/jiwo-0.8b.
| Benchmark | Result |
|---|---|
| Decision Index 0.2.1 | 44.97, 1st of 18 in the 3–6B class on the live board of 2026-09-28, and 2nd of 27 with open submissions¹ |
| Area skill ×100: knowledge / language / retrieval and classification / tools / arts | 26.2 / 49.1 / 51.6 / 67.6 / 28.2 |
| Untrained Qwen3.5-4B, same server and suite, temperatures fitted on the same calibration data | 29.20: +15.8 points for jiwo-4b |
| Best other Qwen3.5-4B fine-tune on the live board | JPT-4B, 43.04: +4.5% for jiwo-4b |
| Best other Qwen3.5-4B fine-tune, including open submissions¹ | ezjev-4b-s2, 51.15: −12.1% for jiwo-4b |
| Median request latency on one H100 | 53 ms |
The score comes from a complete run of the suite (150,317 requests), served by jiwo serve with
JIWO_MAX_LENGTH=65536. The board maintainers did not verify it yet.
¹ Complete runs in open pull requests (as of 2026-10-04) with a readable scores.json and a median latency under 1,000 ms, the limit of the board. One slower Qwen3.5-4B submission also scores higher: Wald-Q4B v1.1 (54.59, median 2,360 ms).
pip install "jiwo @ git+https://github.com/jiwidi/jiwo" # on NVIDIA GPUs: "jiwo[cuda] @ git+..."
from jiwo.model import DecisionModel
model = DecisionModel.from_pretrained("eljiwo/jiwo-4b")
question = {"type": "noul", "instructions": "The customer is angry."}
response = model.decide("Refund please, the parcel arrived crushed.", {"angry": question})
print(response["answers"]["angry"]["noul"]) # the probability that the statement is true
Or serve it over HTTP. The server listens on 127.0.0.1:8765:
JIWO_CHECKPOINT=eljiwo/jiwo-4b jiwo serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"state": "Refund please, the parcel arrived crushed.",
"questions": {
"team": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "Payments and refunds", "shipping": "Damaged or lost parcels", "account": null}},
"angry": {"type": "noul", "instructions": "The customer is angry."},
"urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Not urgent", "Soon", "Now"]}}}'
The jiwo README describes the response, the errors and all settings.
A request has a state (a string, a JSON object or a JSON array) and at most 64 named questions.
| Type | Meaning | criteria |
|---|---|---|
choice | Pick one of N named options | An object that maps option keys to descriptions or null |
score | Rate the state on 2 to 10 ordered levels | A list of level descriptions, lowest level first |
noul | Decide whether a statement is true | Optional. Descriptions for the keys true and false |
The model was fine-tuned on the train splits of public datasets and on synthetic data. The training data is mostly English.
JIWO_MAX_LENGTH, but the model did not see longer prompts in training.choice question can have at most 255 options.Apache-2.0, the licence of the base model. See LICENSE.