Jeeves – Reasoning improves Jev-like decision models
Python
73
20 commits
updated Sep 29, 2026
A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO.
Inspired by Kev.
noul), multiple-choice (choice), and rating (score) questions in the same request, through a Jev-compatible API.Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides.
This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public).
Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.
| bench | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain and held-out, item-weighted) | 0.822 | 0.857 | 0.889 |
| Transfer overall (MMLU-Pro and buried state) | 0.579 | 0.800 | 0.746 |
| JevBench overall (231 public items) | 0.715* | 0.866 | 0.935 |
| QNLI | 0.925 | 0.925 | 0.913 |
| SciQ | 0.963 | 0.988 | 0.991 |
| TweetEval offensive | 0.775 | 0.813 | 0.813 |
| PAWS | 0.763 | 0.788 | 0.875 |
| MMLU | 0.738 | 0.900 | 0.793 |
| Emotion | 0.600 | 0.588 | 0.647 |
| Held-out rule structures | 0.896 | 0.885 | 1.000 |
| Contrastive policies | 0.900 | 0.963 | 1.000 |
| MMLU-Pro (10-way) | 0.515 | 0.840 | 0.739 |
| Buried state | 0.740 | 0.700 | 0.759 |
| Unknowable answered at p ≥ 0.9 (lower is better) | 0.000 | 0.090 | 0.055 |
| JevBench hard (111 public items) | 0.451* | 0.730 | 0.865 |
| JevBench ECE (public items) | 0.049 | 0.037 |
* No Kev-9B JevBench result is published. These are Kev-8B (Qwen3).
All JevBench numbers are on the public easy, standard and hard tiers (231 items). The sealed judge tier is not included, and the Jev and Kev numbers are restricted to the same public items.
Without thinking the same checkpoint scores 0.804 on our test split (2,962 items), against 0.840 with it.
Requirements: Python 3.12 and a CUDA GPU.
pip install -r requirements.txt
Download the released weights and serve them:
hf download PostHog/jeeves --local-dir jeeves-weights
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009
Or fuse your own trained checkpoint into a standalone model and serve it with a drafter:
python export.py runs/cispo/final --out runs/fused
python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009
Then send a request in Jev's format:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
},
"options": {"max_think": 512}}'
Response on one H100 (FP8), with the three questions thinking in parallel:
{
"model": "jeeves-latest",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"confidence": 0.19,
"probabilities": { "returns": 0.4, "shipping": 0.14, "billing": 0.46 }
},
"escalate": { "type": "noul", "noul": 0.72 },
"frustration": {
"type": "score",
"score": 1.5,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.04, "1": 0.43, "2": 0.54 },
"confidence": 0.75
}
},
"usage": { "input_tokens": 129, "output_tokens": 160, "reasoning_tokens": 1536 },
"latency_ms": 8141.6
}
sdk/ is a drop-in replacement for Jev's Python SDK (typesafe-sdk):
pip install ./sdk
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
result = client.system_one(
state="I was charged twice. Please help.",
questions={
"billing": Noul(instructions="Is this about billing?"),
"tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}),
"urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]),
},
max_think=768,
return_reasoning=True,
)
print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score)
print(result.reasoning["tone"].text)
The client connects to http://127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key, and waits up to 120s.
options is optional and ignored by Jev clients that don't send it. Server-wide defaults are set with the matching serve flags.
| option | default | effect |
|---|---|---|
think | true | false answers from the prompt alone (about 0.3 s) |
max_think | 2560 | truncates each reasoning chain at this many tokens, then answers |
nothink_threshold | null | answers without thinking when the no-think confidence is at least this value |
return_reasoning | false | adds each question's reasoning text to the response |
On 325 dev questions:
| setting | accuracy | mean reasoning tokens | median / p90 latency |
|---|---|---|---|
| full thinking | 0.825 | 1,138 | 3.3 s / 17.1 s |
max_think 768, nothink_threshold 0.9 | 0.806 | 344 | 2.0 s / 5.6 s |
| no thinking | 0.775 | 0 | about 0.3 s |
Questions, states and answers are loaded into the Qwen chat template like
<state> …state…
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> …
<think>
The model then rolls out its reasoning chain, and after the </think> token we append
</think>
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> …
<decide>
A pointer head scores each option with a scaled dot product between a query projection of the hidden state at <decide> and a key projection of the hidden state at that option's </opt>, where
<state>, <q>, <opt>, </opt>, <decide> = "<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>"
These are rare, largely unused tokens in the Qwen tokenizer. Ablations found that using plain text like "State" in the prompt instead worsened performance. Likewise, not repeating the questions after the reasoning block also decreases performance. The final probabilities are a softmax over the option scores, divided by a temperature fitted on the dev set.
Stopping at step 402 keeps the best calibration and dev score. Past it, the head over-sharpens on the saturated RL pool.
A diffusion view of the frozen model (drafter/), inspired by Orthrus.
Unlike Orthrus, which supports attention-only models, it supports Qwen3.5's Gated DeltaNet layers by letting mask tokens cross-attend to those layers' post-convolution keys and values.
| chain tokens per second | |
|---|---|
| plain graphed greedy decoding, one question | 109 |
| block 4, one question | 176 (1.6×) |
| block 8, one question | 193 (1.76×) |
| block 4, eight questions batched | about 960 in total |
Block 4 is the default because it stays cheap when several questions are batched.
You can build the datasets locally using the prep scripts. This downloads the public datasets from Hugging Face at the revisions pinned in prep/public.py:
python -m prep.prep
Each public dataset stays under its own license.
On 8 GPUs, with the data in data/, bash run.sh runs the whole pipeline:
torchrun --nproc_per_node 8 train.py sft --run-dir runs/sft
torchrun --nproc_per_node 8 train.py cispo --run-dir runs/cispo --init runs/sft/final
torchrun --nproc_per_node 8 test.py runs/cispo/final
torchrun --nproc_per_node 8 jevbench.py runs/cispo/final
python export.py runs/cispo/final --out runs/fused
torchrun --nproc_per_node 8 -m drafter.gen --model runs/fused
torchrun --nproc_per_node 8 train.py drafter --model runs/fused --block 4 --run-dir runs/drafter_k4
| path | contents |
|---|---|
model/ | Qwen3.5 (Gated DeltaNet + gated attention), LoRA, pointer head |
loader/ | prompt format, tokenisation and batching |
prep/ | dataset construction (prep.py) and synthetic generators |
trainer.py, train.py | SFT, CISPO and drafter training |
test.py, jevbench.py, calibrate.py | evaluation, JevBench, temperature fitting |
export.py | fuses LoRA into a standalone model with the head and temperature |
drafter/ | drafter model, chain sampling, fused speculative decoder |
inference/ | FP8 kernel, batched speculative engine, Jev-compatible server and benchmark |
sdk/ | jeeves_sdk, a drop-in replacement for Jev's Python SDK with the reasoning options |
max_think and nothink_threshold when latency matters.If you use Jeeves, its training recipe or its drafter, please cite:
@software{waltz2026jeeves,
author = {Waltz, Nicholas P.},
title = {Jeeves: Reasoning Improves Jev-like Decisions},
year = {2026},
url = {https://github.com/PostHog/jeeves},
note = {Qwen3.5-9B decision model trained with SFT and CISPO, with a block-4 diffusion drafter}
}
Python
99.9%
Jeeves – Reasoning improves Jev-like decision models
Python
73
20 commits
updated Sep 29, 2026
A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO.
Inspired by Kev.
noul), multiple-choice (choice), and rating (score) questions in the same request, through a Jev-compatible API.Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides.
This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public).
Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.
| bench | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain and held-out, item-weighted) | 0.822 | 0.857 | 0.889 |
| Transfer overall (MMLU-Pro and buried state) | 0.579 | 0.800 | 0.746 |
| JevBench overall (231 public items) | 0.715* | 0.866 | 0.935 |
| QNLI | 0.925 | 0.925 | 0.913 |
| SciQ | 0.963 | 0.988 | 0.991 |
| TweetEval offensive | 0.775 | 0.813 | 0.813 |
| PAWS | 0.763 | 0.788 | 0.875 |
| MMLU | 0.738 | 0.900 | 0.793 |
| Emotion | 0.600 | 0.588 | 0.647 |
| Held-out rule structures | 0.896 | 0.885 | 1.000 |
| Contrastive policies | 0.900 | 0.963 | 1.000 |
| MMLU-Pro (10-way) | 0.515 | 0.840 | 0.739 |
| Buried state | 0.740 | 0.700 | 0.759 |
| Unknowable answered at p ≥ 0.9 (lower is better) | 0.000 | 0.090 | 0.055 |
| JevBench hard (111 public items) | 0.451* | 0.730 | 0.865 |
| JevBench ECE (public items) | 0.049 | 0.037 |
* No Kev-9B JevBench result is published. These are Kev-8B (Qwen3).
All JevBench numbers are on the public easy, standard and hard tiers (231 items). The sealed judge tier is not included, and the Jev and Kev numbers are restricted to the same public items.
Without thinking the same checkpoint scores 0.804 on our test split (2,962 items), against 0.840 with it.
Requirements: Python 3.12 and a CUDA GPU.
pip install -r requirements.txt
Download the released weights and serve them:
hf download PostHog/jeeves --local-dir jeeves-weights
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009
Or fuse your own trained checkpoint into a standalone model and serve it with a drafter:
python export.py runs/cispo/final --out runs/fused
python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009
Then send a request in Jev's format:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
},
"options": {"max_think": 512}}'
Response on one H100 (FP8), with the three questions thinking in parallel:
{
"model": "jeeves-latest",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"confidence": 0.19,
"probabilities": { "returns": 0.4, "shipping": 0.14, "billing": 0.46 }
},
"escalate": { "type": "noul", "noul": 0.72 },
"frustration": {
"type": "score",
"score": 1.5,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.04, "1": 0.43, "2": 0.54 },
"confidence": 0.75
}
},
"usage": { "input_tokens": 129, "output_tokens": 160, "reasoning_tokens": 1536 },
"latency_ms": 8141.6
}
sdk/ is a drop-in replacement for Jev's Python SDK (typesafe-sdk):
pip install ./sdk
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
result = client.system_one(
state="I was charged twice. Please help.",
questions={
"billing": Noul(instructions="Is this about billing?"),
"tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}),
"urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]),
},
max_think=768,
return_reasoning=True,
)
print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score)
print(result.reasoning["tone"].text)
The client connects to http://127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key, and waits up to 120s.
options is optional and ignored by Jev clients that don't send it. Server-wide defaults are set with the matching serve flags.
| option | default | effect |
|---|---|---|
think | true | false answers from the prompt alone (about 0.3 s) |
max_think | 2560 | truncates each reasoning chain at this many tokens, then answers |
nothink_threshold | null | answers without thinking when the no-think confidence is at least this value |
return_reasoning | false | adds each question's reasoning text to the response |
On 325 dev questions:
| setting | accuracy | mean reasoning tokens | median / p90 latency |
|---|---|---|---|
| full thinking | 0.825 | 1,138 | 3.3 s / 17.1 s |
max_think 768, nothink_threshold 0.9 | 0.806 | 344 | 2.0 s / 5.6 s |
| no thinking | 0.775 | 0 | about 0.3 s |
Questions, states and answers are loaded into the Qwen chat template like
<state> …state…
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> …
<think>
The model then rolls out its reasoning chain, and after the </think> token we append
</think>
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> …
<decide>
A pointer head scores each option with a scaled dot product between a query projection of the hidden state at <decide> and a key projection of the hidden state at that option's </opt>, where
<state>, <q>, <opt>, </opt>, <decide> = "<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>"
These are rare, largely unused tokens in the Qwen tokenizer. Ablations found that using plain text like "State" in the prompt instead worsened performance. Likewise, not repeating the questions after the reasoning block also decreases performance. The final probabilities are a softmax over the option scores, divided by a temperature fitted on the dev set.
Stopping at step 402 keeps the best calibration and dev score. Past it, the head over-sharpens on the saturated RL pool.
A diffusion view of the frozen model (drafter/), inspired by Orthrus.
Unlike Orthrus, which supports attention-only models, it supports Qwen3.5's Gated DeltaNet layers by letting mask tokens cross-attend to those layers' post-convolution keys and values.
| chain tokens per second | |
|---|---|
| plain graphed greedy decoding, one question | 109 |
| block 4, one question | 176 (1.6×) |
| block 8, one question | 193 (1.76×) |
| block 4, eight questions batched | about 960 in total |
Block 4 is the default because it stays cheap when several questions are batched.
You can build the datasets locally using the prep scripts. This downloads the public datasets from Hugging Face at the revisions pinned in prep/public.py:
python -m prep.prep
Each public dataset stays under its own license.
On 8 GPUs, with the data in data/, bash run.sh runs the whole pipeline:
torchrun --nproc_per_node 8 train.py sft --run-dir runs/sft
torchrun --nproc_per_node 8 train.py cispo --run-dir runs/cispo --init runs/sft/final
torchrun --nproc_per_node 8 test.py runs/cispo/final
torchrun --nproc_per_node 8 jevbench.py runs/cispo/final
python export.py runs/cispo/final --out runs/fused
torchrun --nproc_per_node 8 -m drafter.gen --model runs/fused
torchrun --nproc_per_node 8 train.py drafter --model runs/fused --block 4 --run-dir runs/drafter_k4
| path | contents |
|---|---|
model/ | Qwen3.5 (Gated DeltaNet + gated attention), LoRA, pointer head |
loader/ | prompt format, tokenisation and batching |
prep/ | dataset construction (prep.py) and synthetic generators |
trainer.py, train.py | SFT, CISPO and drafter training |
test.py, jevbench.py, calibrate.py | evaluation, JevBench, temperature fitting |
export.py | fuses LoRA into a standalone model with the head and temperature |
drafter/ | drafter model, chain sampling, fused speculative decoder |
inference/ | FP8 kernel, batched speculative engine, Jev-compatible server and benchmark |
sdk/ | jeeves_sdk, a drop-in replacement for Jev's Python SDK with the reasoning options |
max_think and nothink_threshold when latency matters.If you use Jeeves, its training recipe or its drafter, please cite:
@software{waltz2026jeeves,
author = {Waltz, Nicholas P.},
title = {Jeeves: Reasoning Improves Jev-like Decisions},
year = {2026},
url = {https://github.com/PostHog/jeeves},
note = {Qwen3.5-9B decision model trained with SFT and CISPO, with a block-4 diffusion drafter}
}
Python
99.9%