Model card: kev-0.5b
9
3 commits
1 linked in READMEs
updated Sep 18, 2026
kev-0.5b is a decision model. It takes one document (the state) and a set of typed questions, and returns a probability distribution for each question in one forward pass. It does not generate text.
It is a LoRA adapter plus a small pointer head on top of Qwen/Qwen2.5-0.5B. It reproduces the architecture that Archer Hume inferred for TypeSafe's Jev in Jev's Architecture Unmasked, and it serves TypeSafe's public /v1/systemone API contract.
This checkpoint is a research prototype trained on a laptop. It shows that the mechanism works. It is not a production model and it is not Jev.
kev)v0.1.0, kev-0.5b.tar.gz (38 MB; LoRA adapter adapter_model.safetensors, head head.pt, tokenizer files, eval.json, training log). SHA-256 15639f79β¦6e12f8, full digest in the sidecar .sha256. Extract to runs/kev/. Weights are not committed to git.| Developed by | Jared Palmer, with Devin (Cognition) |
| Model type | Causal transformer, prefill-only, block-causal branch mask, pointer readout |
| Base model | Qwen/Qwen2.5-0.5B (494M parameters, frozen) |
| Adapter | LoRA rank 16, alpha 32, dropout 0.05, on q_proj k_proj v_proj o_proj gate_proj up_proj down_proj (all 24 layers) |
| Head | Two linear maps 896 β 256 (query from <decide>, key from each </opt>), scaled dot product, softmax over options |
| Trainable parameters | 9.3M (LoRA 8.8M + head 0.46M), 1.9% of the backbone |
| Precision | fp32 (training and serving on Apple MPS) |
| Context used in training | β€ 384 state tokens, β€ 1,024 tokens per question branch |
| Context allowed at serving | 8,192 per branch (backbone supports 32k) |
| Question types | noul (yes/no), choice (2β255 options), score (2β255 ordered levels) |
| Language | English |
| License | Apache-2.0 for the adapter and head. The base model is under the Qwen license (Apache-2.0 for Qwen2.5-0.5B). Datasets carry their own licenses. |
| Version | kev-0.5b v0.1, trained 2026-09-17 |
Intended. Research on decision models: calibration of direct probability readouts, shared-state / isolated-question attention, option-order sensitivity, and API-level compatibility with TypeSafe's System One contract. Local demos and teaching.
Not intended. Any production decision that affects people: moderation, fraud, credit, hiring, medical or legal routing. The model's knowledge is limited to a 0.5B backbone, its calibration is only verified on the training distributions, and its outputs on unfamiliar tasks have not been measured.
Input is one packed token sequence:
<state> β¦stateβ¦ <q> instr <opt> o1 </opt> <opt> o2 </opt> β¦ <decide> <q> β¦ <decide> β¦
</opt> hidden state against the <decide> hidden state and applies softmax.choice/confidence for Choice, p(yes) for Noul, expected level for Score.Reserved tokens are existing Qwen special tokens (<|fim_prefix|>, <|fim_middle|>, <|box_start|>, <|box_end|>, <|fim_suffix|>). User text is sanitized so it cannot produce them.
Serve with python -m kev.serve --run runs/kev and call POST /v1/systemone, or use typesafe-sdk with base_url="http://127.0.0.1:8009".
Six public datasets, converted to TypeSafe-shaped requests and rendered with the same code path used at serving time (api.to_record()). 1,500 records were sampled per source from the standard train splits, giving 9,000 records and 13,500 questions (4,500 Choice, 6,000 Noul, 3,000 Score).
| source | split | converted to | notes |
|---|---|---|---|
| Banking77 | train | Choice, K = 77 | intent names as option keys; templated descriptions, 50% null |
| BoolQ | train | Noul | passage as state; 40% with true/false criteria |
| AG News | train | Choice K = 4 + 2 Noul | derived yes/no questions packed with the topic question |
| MNLI | train | Choice K = 3 | premise as state, hypothesis in instructions |
| SST-5 | train | Score, 5 levels | |
| Yelp Review Full | train | Score 5 levels + Noul | text truncated to 220 words; recommend = stars β₯ 4 |
Rendering variation applied at conversion time: ~30% null option descriptions, ~10% structured {"what": β¦} descriptions, ~15% structured {"question", "focus"} instructions, ~32% states wrapped as objects or arrays ({"document"}, {"ticket": {"channel","body"}}, [{"role","content"}]).
Augmentation applied once per record before encoding: option order shuffled; with probability 0.10 the true option replaced by other: None of the above; with probability 0.15 an irrelevant distractor option added.
No LLM-generated data. No human annotation beyond the original datasets.
| Objective | Cross-entropy over options, averaged over questions in a record |
| Optimizer | AdamW, lr 2e-4, weight decay 0.01, OneCycle schedule (10% warm-up) |
| Batch | 1 record per step, gradient accumulation 8, gradient clipping 1.0 |
| Epochs | 2 (2,250 optimizer steps) |
| Hardware | Apple M5, 32 GB unified memory, PyTorch 2.8 MPS backend |
| Wall clock | ~1h45m (~0.29 s per record) |
| Seed | 0 |
| Final training loss | 0.27 |
This checkpoint predates two loss terms that are now defaults in kev/train.py: the ordinal term for Score (--ord_w) and the permutation-consistency KL for Choice (--perm_kl). To reproduce this checkpoint exactly:
uv run python -m kev.train --n_per_source 1500 --epochs 2 --accum 8 --perm_kl 0 --ord_w 0 --out runs/kev
Note that augmentation is now re-applied every epoch rather than fixed at encode time, so a re-run will not be bit-identical.
Held-out test / validation splits of the same six sources, 150 records per source, 1,350 questions, seed 1. Full results in runs/kev/eval.json.
| source | K | zero-shot base | zero-shot Instruct | kev-0.5b |
|---|---|---|---|---|
| acc / ECE | acc / ECE | acc / ECE / NLL | ||
| banking77 | 77 | β | β | 0.860 / 0.057 / 0.56 |
| agnews | 4 | 0.813 / 0.069 | 0.787 / 0.160 | 0.940 / 0.028 / 0.22 |
| agnews yes/no | 2 | 0.780 / 0.103 | 0.853 / 0.062 | 0.960 / 0.017 / 0.10 |
| boolq | 2 | 0.427 / 0.274 | 0.607 / 0.084 | 0.753 / 0.136 / 0.63 |
| mnli | 3 | 0.460 / 0.225 | 0.433 / 0.390 | 0.747 / 0.100 / 0.63 |
| sst5 | 5 | 0.373 / 0.083 | 0.447 / 0.344 | 0.533 / 0.121 / 1.17 (MAE 0.59 levels) |
| yelp | 5 | 0.313 / 0.043 | 0.353 / 0.078 | 0.553 / 0.118 / 0.95 (MAE 0.54 levels) |
| yelp yes/no | 2 | 0.833 / 0.129 | 0.833 / 0.066 | 0.887 / 0.084 / 0.33 |
| all | 0.799 / 0.065 |
Baselines: Qwen/Qwen2.5-0.5B (raw) and Qwen/Qwen2.5-0.5B-Instruct (chat template), same rendered text, next-token logits over option letters AβH; not run for K = 77. ECE uses 10 equal-width bins on the top probability.
Fit on even-indexed records, tested on odd-indexed: T = 1.47. Held-out NLL 0.505 β 0.481, ECE 0.057 β 0.031. The model is mildly over-confident before scaling.
| test | result |
|---|---|
| Isolation (secret in sibling question / absent / in state) | p = 0.03 / 0.03 / 0.99 |
| Packed vs separate, max abs probability difference | 3.7e-6 (2.0Γ faster packed, ~2.7 questions per request) |
| Permutation, 4 orders, Choice K β₯ 3 | argmax flips 7.4%; mean spread of p(correct) 0.065, p90 0.25 |
| IIA, append one irrelevant option | mean |Ξ log-odds| top-2 = 0.13, p90 0.34 |
| Boundary forgery, option text with fake delimiters | option count unchanged; forged option p β€ 0.09 |
return_policy where Jev picks return_status. Reading comprehension (BoolQ 0.75, MNLI 0.75) is far below state of the art.1 β E|level β mode| / (L β 1); TypeSafe's formula is unpublished.The training sets carry the biases of their sources: US-centric news categories, English banking terminology, restaurant reviews, and crowd-sourced NLI labels. The model will mirror them.
Direct probability outputs look authoritative. A confidence: 0.92 from this model is a statistic about its own distribution over three options, not a verified probability of being right. Do not threshold on it for consequential decisions without measuring calibration on your own labelled outcomes first.
The question-isolation property is a real safety feature (one question's text cannot manipulate another's answer) and was verified. The delimiter-forgery protection was verified for the five reserved tokens. Other prompt-injection routes through the state text have not been studied.
One training run: ~1.75 h on a single Apple M5 laptop SoC at roughly 30β40 W, i.e. about 0.06 kWh. Evaluation and smoke runs add a similar amount. This is small.
@software{kev2026,
title = {kev: a laptop-scale reconstruction of a Jev-style decision model},
author = {Palmer, Jared},
year = {2026},
url = {https://github.com/jaredpalmer/kev}
}
@misc{hume2026jev,
title = {Jev's Architecture Unmasked},
author = {Hume, Archer},
year = {2026},
url = {https://archerhume.com/posts/jevs-architecture-unmasked}
}
Open an issue at github.com/jaredpalmer/kev.
3 commits
Model card: kev-0.5b
9
3 commits
1 linked in READMEs
updated Sep 18, 2026
kev-0.5b is a decision model. It takes one document (the state) and a set of typed questions, and returns a probability distribution for each question in one forward pass. It does not generate text.
It is a LoRA adapter plus a small pointer head on top of Qwen/Qwen2.5-0.5B. It reproduces the architecture that Archer Hume inferred for TypeSafe's Jev in Jev's Architecture Unmasked, and it serves TypeSafe's public /v1/systemone API contract.
This checkpoint is a research prototype trained on a laptop. It shows that the mechanism works. It is not a production model and it is not Jev.
kev)v0.1.0, kev-0.5b.tar.gz (38 MB; LoRA adapter adapter_model.safetensors, head head.pt, tokenizer files, eval.json, training log). SHA-256 15639f79β¦6e12f8, full digest in the sidecar .sha256. Extract to runs/kev/. Weights are not committed to git.| Developed by | Jared Palmer, with Devin (Cognition) |
| Model type | Causal transformer, prefill-only, block-causal branch mask, pointer readout |
| Base model | Qwen/Qwen2.5-0.5B (494M parameters, frozen) |
| Adapter | LoRA rank 16, alpha 32, dropout 0.05, on q_proj k_proj v_proj o_proj gate_proj up_proj down_proj (all 24 layers) |
| Head | Two linear maps 896 β 256 (query from <decide>, key from each </opt>), scaled dot product, softmax over options |
| Trainable parameters | 9.3M (LoRA 8.8M + head 0.46M), 1.9% of the backbone |
| Precision | fp32 (training and serving on Apple MPS) |
| Context used in training | β€ 384 state tokens, β€ 1,024 tokens per question branch |
| Context allowed at serving | 8,192 per branch (backbone supports 32k) |
| Question types | noul (yes/no), choice (2β255 options), score (2β255 ordered levels) |
| Language | English |
| License | Apache-2.0 for the adapter and head. The base model is under the Qwen license (Apache-2.0 for Qwen2.5-0.5B). Datasets carry their own licenses. |
| Version | kev-0.5b v0.1, trained 2026-09-17 |
Intended. Research on decision models: calibration of direct probability readouts, shared-state / isolated-question attention, option-order sensitivity, and API-level compatibility with TypeSafe's System One contract. Local demos and teaching.
Not intended. Any production decision that affects people: moderation, fraud, credit, hiring, medical or legal routing. The model's knowledge is limited to a 0.5B backbone, its calibration is only verified on the training distributions, and its outputs on unfamiliar tasks have not been measured.
Input is one packed token sequence:
<state> β¦stateβ¦ <q> instr <opt> o1 </opt> <opt> o2 </opt> β¦ <decide> <q> β¦ <decide> β¦
</opt> hidden state against the <decide> hidden state and applies softmax.choice/confidence for Choice, p(yes) for Noul, expected level for Score.Reserved tokens are existing Qwen special tokens (<|fim_prefix|>, <|fim_middle|>, <|box_start|>, <|box_end|>, <|fim_suffix|>). User text is sanitized so it cannot produce them.
Serve with python -m kev.serve --run runs/kev and call POST /v1/systemone, or use typesafe-sdk with base_url="http://127.0.0.1:8009".
Six public datasets, converted to TypeSafe-shaped requests and rendered with the same code path used at serving time (api.to_record()). 1,500 records were sampled per source from the standard train splits, giving 9,000 records and 13,500 questions (4,500 Choice, 6,000 Noul, 3,000 Score).
| source | split | converted to | notes |
|---|---|---|---|
| Banking77 | train | Choice, K = 77 | intent names as option keys; templated descriptions, 50% null |
| BoolQ | train | Noul | passage as state; 40% with true/false criteria |
| AG News | train | Choice K = 4 + 2 Noul | derived yes/no questions packed with the topic question |
| MNLI | train | Choice K = 3 | premise as state, hypothesis in instructions |
| SST-5 | train | Score, 5 levels | |
| Yelp Review Full | train | Score 5 levels + Noul | text truncated to 220 words; recommend = stars β₯ 4 |
Rendering variation applied at conversion time: ~30% null option descriptions, ~10% structured {"what": β¦} descriptions, ~15% structured {"question", "focus"} instructions, ~32% states wrapped as objects or arrays ({"document"}, {"ticket": {"channel","body"}}, [{"role","content"}]).
Augmentation applied once per record before encoding: option order shuffled; with probability 0.10 the true option replaced by other: None of the above; with probability 0.15 an irrelevant distractor option added.
No LLM-generated data. No human annotation beyond the original datasets.
| Objective | Cross-entropy over options, averaged over questions in a record |
| Optimizer | AdamW, lr 2e-4, weight decay 0.01, OneCycle schedule (10% warm-up) |
| Batch | 1 record per step, gradient accumulation 8, gradient clipping 1.0 |
| Epochs | 2 (2,250 optimizer steps) |
| Hardware | Apple M5, 32 GB unified memory, PyTorch 2.8 MPS backend |
| Wall clock | ~1h45m (~0.29 s per record) |
| Seed | 0 |
| Final training loss | 0.27 |
This checkpoint predates two loss terms that are now defaults in kev/train.py: the ordinal term for Score (--ord_w) and the permutation-consistency KL for Choice (--perm_kl). To reproduce this checkpoint exactly:
uv run python -m kev.train --n_per_source 1500 --epochs 2 --accum 8 --perm_kl 0 --ord_w 0 --out runs/kev
Note that augmentation is now re-applied every epoch rather than fixed at encode time, so a re-run will not be bit-identical.
Held-out test / validation splits of the same six sources, 150 records per source, 1,350 questions, seed 1. Full results in runs/kev/eval.json.
| source | K | zero-shot base | zero-shot Instruct | kev-0.5b |
|---|---|---|---|---|
| acc / ECE | acc / ECE | acc / ECE / NLL | ||
| banking77 | 77 | β | β | 0.860 / 0.057 / 0.56 |
| agnews | 4 | 0.813 / 0.069 | 0.787 / 0.160 | 0.940 / 0.028 / 0.22 |
| agnews yes/no | 2 | 0.780 / 0.103 | 0.853 / 0.062 | 0.960 / 0.017 / 0.10 |
| boolq | 2 | 0.427 / 0.274 | 0.607 / 0.084 | 0.753 / 0.136 / 0.63 |
| mnli | 3 | 0.460 / 0.225 | 0.433 / 0.390 | 0.747 / 0.100 / 0.63 |
| sst5 | 5 | 0.373 / 0.083 | 0.447 / 0.344 | 0.533 / 0.121 / 1.17 (MAE 0.59 levels) |
| yelp | 5 | 0.313 / 0.043 | 0.353 / 0.078 | 0.553 / 0.118 / 0.95 (MAE 0.54 levels) |
| yelp yes/no | 2 | 0.833 / 0.129 | 0.833 / 0.066 | 0.887 / 0.084 / 0.33 |
| all | 0.799 / 0.065 |
Baselines: Qwen/Qwen2.5-0.5B (raw) and Qwen/Qwen2.5-0.5B-Instruct (chat template), same rendered text, next-token logits over option letters AβH; not run for K = 77. ECE uses 10 equal-width bins on the top probability.
Fit on even-indexed records, tested on odd-indexed: T = 1.47. Held-out NLL 0.505 β 0.481, ECE 0.057 β 0.031. The model is mildly over-confident before scaling.
| test | result |
|---|---|
| Isolation (secret in sibling question / absent / in state) | p = 0.03 / 0.03 / 0.99 |
| Packed vs separate, max abs probability difference | 3.7e-6 (2.0Γ faster packed, ~2.7 questions per request) |
| Permutation, 4 orders, Choice K β₯ 3 | argmax flips 7.4%; mean spread of p(correct) 0.065, p90 0.25 |
| IIA, append one irrelevant option | mean |Ξ log-odds| top-2 = 0.13, p90 0.34 |
| Boundary forgery, option text with fake delimiters | option count unchanged; forged option p β€ 0.09 |
return_policy where Jev picks return_status. Reading comprehension (BoolQ 0.75, MNLI 0.75) is far below state of the art.1 β E|level β mode| / (L β 1); TypeSafe's formula is unpublished.The training sets carry the biases of their sources: US-centric news categories, English banking terminology, restaurant reviews, and crowd-sourced NLI labels. The model will mirror them.
Direct probability outputs look authoritative. A confidence: 0.92 from this model is a statistic about its own distribution over three options, not a verified probability of being right. Do not threshold on it for consequential decisions without measuring calibration on your own labelled outcomes first.
The question-isolation property is a real safety feature (one question's text cannot manipulate another's answer) and was verified. The delimiter-forgery protection was verified for the five reserved tokens. Other prompt-injection routes through the state text have not been studied.
One training run: ~1.75 h on a single Apple M5 laptop SoC at roughly 30β40 W, i.e. about 0.06 kWh. Evaluation and smoke runs add a similar amount. This is small.
@software{kev2026,
title = {kev: a laptop-scale reconstruction of a Jev-style decision model},
author = {Palmer, Jared},
year = {2026},
url = {https://github.com/jaredpalmer/kev}
}
@misc{hume2026jev,
title = {Jev's Architecture Unmasked},
author = {Hume, Archer},
year = {2026},
url = {https://archerhume.com/posts/jevs-architecture-unmasked}
}
Open an issue at github.com/jaredpalmer/kev.
3 commits