eljiwo/jiwo-4b

Model

jiwo-4b

0

4 commits

2 linked in READMEs

updated Oct 5, 2026

See the code

README

jiwo-4b

1st of 18 models in the 3–6B class of the Decision Index live board (44.97), and 2nd of 27 when open submissions are included.

jiwo-4b is a decision model. It reads a state and one or more typed questions. For each question, it returns a probability for every option in one forward pass. It does not generate text.

The request and response format is the format of TypeSafe's Jev API (POST /v1/systemone). jiwo is an independent project. TypeSafe does not endorse it.

Base modelQwen/Qwen3.5-4B
Parameters4.2B
WeightsFull weights: decoder, tokenizer and readout
Question typeschoice (at most 255 options), score (2 to 10 levels), noul (true or false)
Training context length8,192 tokens
LicenceApache-2.0

Code, server and documentation: github.com/jiwidi/jiwo. The other model of the family is eljiwo/jiwo-0.8b.

Evaluation

BenchmarkResult
Decision Index 0.2.144.97, 1st of 18 in the 3–6B class on the live board of 2026-09-28, and 2nd of 27 with open submissions¹
Area skill ×100: knowledge / language / retrieval and classification / tools / arts26.2 / 49.1 / 51.6 / 67.6 / 28.2
Untrained Qwen3.5-4B, same server and suite, temperatures fitted on the same calibration data29.20: +15.8 points for jiwo-4b
Best other Qwen3.5-4B fine-tune on the live boardJPT-4B, 43.04: +4.5% for jiwo-4b
Best other Qwen3.5-4B fine-tune, including open submissions¹ezjev-4b-s2, 51.15: −12.1% for jiwo-4b
Median request latency on one H10053 ms

The score comes from a complete run of the suite (150,317 requests), served by jiwo serve with JIWO_MAX_LENGTH=65536. The board maintainers did not verify it yet.

¹ Complete runs in open pull requests (as of 2026-10-04) with a readable scores.json and a median latency under 1,000 ms, the limit of the board. One slower Qwen3.5-4B submission also scores higher: Wald-Q4B v1.1 (54.59, median 2,360 ms).

Decision Index 0.2.1 by size class

Use the model

pip install "jiwo @ git+https://github.com/jiwidi/jiwo"    # on NVIDIA GPUs: "jiwo[cuda] @ git+..."
from jiwo.model import DecisionModel

model = DecisionModel.from_pretrained("eljiwo/jiwo-4b")
question = {"type": "noul", "instructions": "The customer is angry."}
response = model.decide("Refund please, the parcel arrived crushed.", {"angry": question})
print(response["answers"]["angry"]["noul"])  # the probability that the statement is true

Or serve it over HTTP. The server listens on 127.0.0.1:8765:

JIWO_CHECKPOINT=eljiwo/jiwo-4b jiwo serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Refund please, the parcel arrived crushed.",
  "questions": {
    "team":  {"type": "choice", "instructions": "Which team handles this?",
              "criteria": {"billing": "Payments and refunds", "shipping": "Damaged or lost parcels", "account": null}},
    "angry": {"type": "noul", "instructions": "The customer is angry."},
    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Not urgent", "Soon", "Now"]}}}'

The jiwo README describes the response, the errors and all settings.

Request format

A request has a state (a string, a JSON object or a JSON array) and at most 64 named questions.

TypeMeaningcriteria
choicePick one of N named optionsAn object that maps option keys to descriptions or null
scoreRate the state on 2 to 10 ordered levelsA list of level descriptions, lowest level first
noulDecide whether a statement is trueOptional. Descriptions for the keys true and false

How the model works

  • The decoder of the base model reads one prompt for each (state, question) pair.
  • A linear readout with one row per answer code (A, B, ..., Z, AA, ...) gives one logit for each option.
  • A softmax with a fitted temperature for each question type gives the probabilities.

Training data

The model was fine-tuned on the train splits of public datasets and on synthetic data. The training data is mostly English.

Limits

  • The training prompts had at most 8,192 tokens. By default, the server refuses a longer prompt with a 400 error. You can set a larger JIWO_MAX_LENGTH, but the model did not see longer prompts in training.
  • A choice question can have at most 255 options.
  • The model was trained mostly on English. Other languages are not evaluated.
  • The model gives probabilities, not explanations. Check its answers before you use them for decisions with a high cost.

Licence

Apache-2.0, the licence of the base model. See LICENSE.

calibrated-probabilities
decision-model
jiwo
qwen3_5_text
safetensors
text-classification

eljiwo/jiwo-4b

Model

jiwo-4b

0

4 commits

2 linked in READMEs

updated Oct 5, 2026

See the code

README

jiwo-4b

1st of 18 models in the 3–6B class of the Decision Index live board (44.97), and 2nd of 27 when open submissions are included.

jiwo-4b is a decision model. It reads a state and one or more typed questions. For each question, it returns a probability for every option in one forward pass. It does not generate text.

The request and response format is the format of TypeSafe's Jev API (POST /v1/systemone). jiwo is an independent project. TypeSafe does not endorse it.

Base modelQwen/Qwen3.5-4B
Parameters4.2B
WeightsFull weights: decoder, tokenizer and readout
Question typeschoice (at most 255 options), score (2 to 10 levels), noul (true or false)
Training context length8,192 tokens
LicenceApache-2.0

Code, server and documentation: github.com/jiwidi/jiwo. The other model of the family is eljiwo/jiwo-0.8b.

Evaluation

BenchmarkResult
Decision Index 0.2.144.97, 1st of 18 in the 3–6B class on the live board of 2026-09-28, and 2nd of 27 with open submissions¹
Area skill ×100: knowledge / language / retrieval and classification / tools / arts26.2 / 49.1 / 51.6 / 67.6 / 28.2
Untrained Qwen3.5-4B, same server and suite, temperatures fitted on the same calibration data29.20: +15.8 points for jiwo-4b
Best other Qwen3.5-4B fine-tune on the live boardJPT-4B, 43.04: +4.5% for jiwo-4b
Best other Qwen3.5-4B fine-tune, including open submissions¹ezjev-4b-s2, 51.15: −12.1% for jiwo-4b
Median request latency on one H10053 ms

The score comes from a complete run of the suite (150,317 requests), served by jiwo serve with JIWO_MAX_LENGTH=65536. The board maintainers did not verify it yet.

¹ Complete runs in open pull requests (as of 2026-10-04) with a readable scores.json and a median latency under 1,000 ms, the limit of the board. One slower Qwen3.5-4B submission also scores higher: Wald-Q4B v1.1 (54.59, median 2,360 ms).

Decision Index 0.2.1 by size class

Use the model

pip install "jiwo @ git+https://github.com/jiwidi/jiwo"    # on NVIDIA GPUs: "jiwo[cuda] @ git+..."
from jiwo.model import DecisionModel

model = DecisionModel.from_pretrained("eljiwo/jiwo-4b")
question = {"type": "noul", "instructions": "The customer is angry."}
response = model.decide("Refund please, the parcel arrived crushed.", {"angry": question})
print(response["answers"]["angry"]["noul"])  # the probability that the statement is true

Or serve it over HTTP. The server listens on 127.0.0.1:8765:

JIWO_CHECKPOINT=eljiwo/jiwo-4b jiwo serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Refund please, the parcel arrived crushed.",
  "questions": {
    "team":  {"type": "choice", "instructions": "Which team handles this?",
              "criteria": {"billing": "Payments and refunds", "shipping": "Damaged or lost parcels", "account": null}},
    "angry": {"type": "noul", "instructions": "The customer is angry."},
    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Not urgent", "Soon", "Now"]}}}'

The jiwo README describes the response, the errors and all settings.

Request format

A request has a state (a string, a JSON object or a JSON array) and at most 64 named questions.

TypeMeaningcriteria
choicePick one of N named optionsAn object that maps option keys to descriptions or null
scoreRate the state on 2 to 10 ordered levelsA list of level descriptions, lowest level first
noulDecide whether a statement is trueOptional. Descriptions for the keys true and false

How the model works

  • The decoder of the base model reads one prompt for each (state, question) pair.
  • A linear readout with one row per answer code (A, B, ..., Z, AA, ...) gives one logit for each option.
  • A softmax with a fitted temperature for each question type gives the probabilities.

Training data

The model was fine-tuned on the train splits of public datasets and on synthetic data. The training data is mostly English.

Limits

  • The training prompts had at most 8,192 tokens. By default, the server refuses a longer prompt with a 400 error. You can set a larger JIWO_MAX_LENGTH, but the model did not see longer prompts in training.
  • A choice question can have at most 255 options.
  • The model was trained mostly on English. Other languages are not evaluated.
  • The model gives probabilities, not explanations. Check its answers before you use them for decisions with a high cost.

Licence

Apache-2.0, the licence of the base model. See LICENSE.

calibrated-probabilities
decision-model
jiwo
qwen3_5_text
safetensors
text-classification