Open Jev-style typed-decision model that also takes images: photo + app state + typed questions in, calibrated probabilities out, locally.
Python
69
14 commits
updated Sep 28, 2026
Small open models (2B · 4B · 9B) that read the photos, records and text a business already has and answer in the options you set,
with a probability on each and an explicit can't tell. Your system acts when it is sure and hands the rest to a person.
Live demo · Website · Quickstart · Checked examples · Results · Technical report
Screenshots of the official leaderboards, captured 28 Sep 2026. Each board is run by its own maintainer, not by us; click through for the live page.
Scored 27 Sep 2026. imajev-4b 67.37, ahead of Plumb-4B (65.84) and TypeSafe's own Jev 1.13.0 (63.29).
Released 28 Sep 2026. imajev-4b 76.39, ahead of Jev-Omni (12B, 73.10) and NeoHorse Jev 4B (71.94); imajev-2b is #6. Sealed accuracy 83.3%, second of 44 self-hosted systems. Measured by the maintainer on the fast server (--fast --merge-lora, one option order, shipped calibration).
"Imajev-4B leads the Jev-class systems with 66.3" (Capability: intelligence and calibration, among systems within 2× Jev's cost and latency).
79.65, ahead of GLM-5.3 Flash (320B), Jev 1.13, DeepSeek V4.1 Flash (552B) and GPT-5.6 Luna. The two above are the benchmark team's own Bosun models.
80.58, behind GLM-5.3 Flash and GPT-5.6 Luna, ahead of DeepSeek V4.1 Flash (552B) and Jev 1.13.
Boards move as new systems are added; ranks are quoted with the board version and the date. Archived copies and the raw data are linked in Official leaderboards.

The live demo runs imajev-4b on a GPU. Pick one of the checked examples (a listing against its photo, a return, a part on the line, an email against a CRM record, a refund against the policy, a ticket, a stylist request it declines to guess on), change the record or swap a photo, and watch the answer and the app's action change. Or upload your own photo and write your own questions.
Send the evidence and the questions you care about. Each answer comes back as a probability over the answers you allowed, ready
for an if. This is the exact script we ran against imajev-4b (site/showcase/listing.py) and its output, with numbers rounded to
three places and usage shortened; 1.15 s on a Mac Studio (four option orders averaged, calibration file applied).
import json, requests
URL = "http://127.0.0.1:8765/v1/systemone"
listing = {
"title": "Men's suede boat shoes",
"color": "red",
"product_type": "shoe",
}
questions = {
"contradicted_field": {
"type": "choice",
"instructions":
"Which field of `listing` does this photo contradict?",
"criteria": {
"listing.color": None,
"listing.product_type": None,
"none of these": "the photo agrees with every field",
},
},
"color_matches": {
"type": "noul",
"instructions":
"The product in the photo matches `listing.color`.",
},
"type_matches": {
"type": "noul",
"instructions": "The photo shows the kind of product "
"given in `listing.product_type`.",
},
}
request = {"state": {"listing": listing}, "questions": questions}
with open("listing.jpg", "rb") as photo:
r = requests.post(URL, files={"image": photo},
data={"request": json.dumps(request)})
print(json.dumps(r.json(), indent=2))
{
"model": "imajev-4b",
"answers": {
"contradicted_field": {
"type": "choice",
"choice": "listing.color",
"probabilities": {
"listing.color": 0.95,
"listing.product_type": 0.006,
"none of these": 0.043
},
"confidence": 0.919,
"unknown_probability": 0.007,
"abstained": false
},
"color_matches": {
"type": "noul",
"noul": 0.082,
"unknown_probability": 0.022,
"abstained": false
},
"type_matches": {
"type": "noul",
"noul": 0.989,
"unknown_probability": 0.004,
"abstained": false
}
},
"usage": {
"total_ms": 1152.9,
"input_tokens": 224
}
}

The same endpoint with no photo: one request routes the ticket (choice), flags urgency (yes/no) and scores frustration (score). Script: site/showcase/ticket.py.
Jev reads text only. General vision models answer in prose, from a hosted API, in seconds. imajev brings Jev's typed answers to photos, and runs on your own hardware.

unknown. Asked for a white or beige bag from a closet with only
a red and a black backpack, the 4B puts 0.89 on unknown instead of guessing.images, unknown_probability and
abstained. Text-only Jev requests work unchanged.| imajev | Jev (TypeSafe) | Jev-Omni | Frontier vision APIs | |
|---|---|---|---|---|
| photos in a request | up to 2 (reference + target) | none, text only | yes; two-photo requests not documented | yes |
| answer format | probability per option you set | probability per option you set | probability per option you set | generated text or JSON |
| says it can't tell | trained unknown, 18 / 21 on ImajevBench (4B) | not documented | no abstain output, so 0 / 21 | only if prompted |
| record per request | 32 KB (about 8k tokens) | 32k tokens | not documented | large |
| where it runs | your hardware, open weights | hosted API | your hardware, open weights (12B) | hosted API |
| time per decision | about 0.1 s raw, about 0.35 s as shipped (one H100) | not documented | about 0.1 s (one H100) | 5 to 8 s |
| ImajevBench accuracy | 83.9% (4B) | cannot take photos | 78.5% | 91.4% to 99.6% |
Jev from docs.typesafe.ai (models page, Jev 1.13.0). Jev-Omni from its model card and our run of its own predict() API.
Frontier rows and all ImajevBench numbers from our runs, 24 Sept 2026; frontier models answer by structured generation, a different
interface. Jev's record limit is larger than imajev's.
The decisions it handles well are the high-volume, well-defined ones with a clear set of answers.
| Area | Decisions (each a checked example on the website) |
|---|---|
| Marketplaces and retail | listing matches its photo · return is the item we shipped · tag a product from a photo |
| Manufacturing and field work | reject a chipped part against a known-good reference · spot what changed on site |
| Customer support | route a ticket and flag urgency · refund against the policy · send a review to the right team |
| Trust and safety | remove spam, harassment and doxxing · ask for a better photo |
| Back office and records | email contradicts the CRM · is the invoice paid? · catalogue record is wrong |
You choose how sure the model must be before it acts. A higher bar automates less and makes fewer mistakes; everything below it goes to a person, including when it says it can't tell.

| Act when at least… | imajev-2b | imajev-4b | imajev-9b |
|---|---|---|---|
| 80% sure | 49% automated, 92.7% right | 63% automated, 94.9% right | 77% automated, 87.9% right |
| 90% sure | 38% automated, 95.3% right | 58% automated, 97.5% right | 70% automated, 91.8% right |
| 99% sure | 21% automated, 100% right | 40% automated, 100% right | 52% automated, 99.3% right |
The 279 ImajevBench test questions (photos, records and text; 21 whose honest answer is can't tell), raw probabilities, scored with the benchmark's own rule (as shipped: four option orders, calibration file). The benchmark is built to be hard, so treat these as a starting point and measure on a few hundred of your own cases before choosing.
Every example comes from one of five small apps in scripts/playground/. Every combination a visitor can click in them is sent to
the model and compared with the right answer (node scripts/playground/verify_scenarios.mjs); a check passes when the answer is
right and the app takes the expected action at an 80% threshold. imajev-4b, served without its calibration file (four option orders):
| App | What it asks | Passed | With the calibration file | |
|---|---|---|---|---|
![]() | Business checks | Does the photo match the listing? Is the return the item we shipped? Is this part chipped? | 19 / 21 | 19 / 21 |
![]() | Text only | An email against a CRM record, a refund against the policy, a post against forum rules, a review, an inbox. Written for the launch and run once. | 20 / 21 | 18 / 21 |
![]() | Wardrobe | Is this what I ordered, does it meet the dress code, do I already own it, which shoes match? | 48 / 49 | 38 / 49 |
![]() | Stylist app | Reads a piece of clothing, then picks bottoms, shoes and a bag from your closet in your colours. | 21 / 24 | 21 / 24 |
![]() | Tracing pad | Reads which letter or number a child traced; the app checks the strokes covered every line. | 22 / 30 | 22 / 30 |
| All | 129 / 145 | 113 / 145 |
With the calibration file the top answer never changes, but confidence is lower, so more cases go to a person at 80%. The misses
worth knowing: with a blank payment note the 4B answered "not paid" instead of unknown, and it reads 8 of 10 scribbles on the tracing
pad as letters. Every run, pass or miss, is in results/scenarios/.
2026-09-26: the 4B moved to its phase-3 adapter (rank-64 LoRA, 256-code readout, trained on the decisions the previous release got wrong): ImajevBench 82.4 → 83.9%, hidden split 84.2 → 85.6%, JevBench hard as shipped 70.3 → 72.1%, DecisionBench full suite 77.5 → 79.7%; it abstains on 18 of 21 ImajevBench Unknown items (was 14) and on 9 of 258 answerable ones (was 3). Its calibration on DecisionBench is worse than before (ECE 0.024 → 0.069). The 2B and 9B are unchanged. Details:
results/phase3/benchmarks.md.
| imajev-2b (latency) | imajev-4b (recommended default) | imajev-9b (quality) | |
|---|---|---|---|
| Base | Qwen3.5-2B (Apache-2.0) | Qwen3.5-4B (Apache-2.0) | Qwen3.5-9B (Apache-2.0) |
| Adapter | LoRA r16/α32 on the language layers + 255-code decision readout; vision tower frozen; shipped as a weight-space average of two adapters (the hard-question adapter and a soft-target continuation of it) | same | same |
| ImajevBench v2.0-lite test | 71.7% | 83.9% | 82.1% |
| JevBench public hard (111), as shipped (4 rotations + calibration) | 60.4% | 72.1% | 69.4% |
| p50 per decision, JevBench hard item, 1×H100, serial, under load | 238 ms as shipped (83 ms raw) | 350 ms (96 ms raw) | 316 ms (96 ms raw) |
| Use it when | latency or memory is the constraint | almost always: within noise of the 9B on ImajevBench | knowledge-heavy text questions, and memory is not a constraint (~19 GB resident) |
| Weights | mohit67890/imajev-2b | mohit67890/imajev-4b | mohit67890/imajev-9b |
Start with the 4B. On ImajevBench it is ahead of the 9B (83.9% vs 82.1%; not significant, paired test p = 0.57) and one point ahead of it on JevBench hard. The 2B is 11 points lower on ImajevBench and 10 lower on JevBench hard; pick it when its footprint is the point. The Mac (MLX) weights agree with the GPU run on 97 to 99% of ImajevBench answers (2B 97.1%, 4B 98.2%, 9B 99.3%).
git clone https://github.com/mohit67890/imajev && cd imajev
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e ".[serve,mlx]" # Apple silicon; elsewhere: pip install -e ".[serve,torch]"
python scripts/download_model.py --model 4b # pinned Qwen3.5-4B into .cache/
hf download mohit67890/imajev-4b --local-dir adapters/imajev-4b
PYTHONPATH=src:scripts python scripts/playground/server.py --model-bundle artifacts/model-qwen4b.json \
--adapter adapters/imajev-4b/mlx --calibration adapters/imajev-4b/calibration.json --model-name imajev-4b --port 8765
Open http://127.0.0.1:8765/ for the playground, or call the API:
curl -s http://127.0.0.1:8765/v1/systemone \
-F 'request={"state":{"listing":{"title":"Blue ceramic mug, 350 ml","colour":"blue"}},
"questions":{"matches":{"type":"noul","instructions":"Does the photo show the listed item?"},
"wrong_field":{"type":"choice","instructions":"Which listing field does the photo contradict?",
"criteria":{"title":null,"colour":null,"none":null}}}}' \
-F image=@photo.jpg
import requests
r = requests.post("http://127.0.0.1:8765/v1/systemone", json={
"state": "Ticket: 'Charged twice for one order, need the duplicate refunded.'",
"questions": {"queue": {"type": "choice", "instructions": "Route the ticket.",
"criteria": {"billing": None, "shipping": None, "account": None, "other": None}},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["can wait a week", "this week", "today", "within the hour", "right now"]}}})
print(r.json()["answers"]["queue"]) # {"type": "choice", "choice": "billing", "probabilities": {...}, "confidence": ..., "unknown_probability": ..., "abstained": false}
On Linux / CUDA, add --backend torch and pass --adapter adapters/imajev-4b (the PEFT adapter at the repo root). For the other
sizes, swap 4b for 2b or 9b in the download commands, the bundle (artifacts/model-qwen9b.json for the 9B; the 2B is the
default bundle) and the adapter paths. --rotations 4 averages four option orders; every number in this README was measured with it and with --calibration (on JevBench hard it adds +1.8 / +0.9 / +0.0 points for the 2B / 4B / 9B at about 3×
the latency). The 9B needs ~19 GB resident; do not keep it and another model loaded on the same Mac.
On a CUDA GPU, --fast makes the torch backend quicker without changing what it computes: one tokenization per question,
image normalisation on the GPU (pixels bit-identical to the processor's), and CUDA graphs of the language model recorded at
load (about a minute; needs a C compiler for the Triton kernels, e.g. build-essential). --merge-lora also folds the adapter
into the weights at load (float32 sum, rounded once to bf16). Checked against a float32 reference on JevBench public and 300
ImajevBench images: --fast --merge-lora is as close to it as the default path, at 11 ms instead of 59 ms per standard text
decision and 91 ms instead of 149 ms per image decision on one H100 (results/serving/fast-path-2026-09-27, scripts/bench_fast_path.py).
What comes back. For each question: choice / score return probabilities over your options (summing to 1, given that the
model answers); noul returns P(yes) with half of the unknown mass added, as in Jev. Every answer also has unknown_probability
(mass on the trained unknown: missing, contradictory or out-of-scope evidence), abstained (unknown was the most likely outcome)
and confidence. Limits per request: up to 2 images, a state up to 32 KB, 1 to 8 questions, 2 to 254 options, 2 to 10 score levels.
Acting on it. Act above a threshold you choose from your own error costs; send abstentions and low-confidence answers to a person.
a = response["answers"]["contradicted_field"]
p = a["probabilities"][a["choice"]] * (1 - a["unknown_probability"])
if a["abstained"] or p < 0.85:
route_to_person(a) # the model cannot tell, or is not sure enough
elif a["choice"] == "none of these":
publish()
else:
hold(field=a["choice"])
Choosing a size. Start with imajev-4b. Use the 2B when latency or memory is tight (it abstains less often than it should on unknown items); use the 9B for knowledge-heavy text questions when ~19 GB of weights is fine.
Checking it on your data. Score a few hundred of your own labelled requests (include "cannot tell" cases) before trusting a
threshold. To fit your own temperature, the evaluators in scripts/ write one JSONL row per question with its logits, and
scripts/v1_text/fit_temperature_calibration.py rows.jsonl calibration.json fits one temperature per question type and option count;
serve it with --calibration. The full guide is section 5 of the technical report.
Measured by each benchmark's maintainer, not by us. Boards move; every rank is quoted with its version and date.
| Board | imajev result | Source |
|---|---|---|
| JevBench v1.4.2.2 (Benchmark Heaven, scored 27 Sep 2026; 91 ranked systems) | imajev-4b #1, JevBench Score 67.37 (Intelligence 52.2, Calibration 80.4, Speed 90.6, Cost 59.7). Plumb-4B 65.84, decider-4b v2 64.13, Jev 1.13.0 63.29. #1 under four of the board's five weightings (#2 under speed-heavy 20:60:20, #3 on Intelligence alone); best Calibration of the top 8. Also #1 on the board's Capability ranking of Jev-class systems (mean of Intelligence and Calibration): 66.3 vs Jev 1.13.0 64.7. Cost on the board's estimate: $0.022 per 1,000 decisions (Jev $0.040). | board · data |
| DecisionBench (eng, v1) (Hanno-Labs, 23 tasks, 23,900 rows; 56 models) | imajev-4b #3, 79.65 (mean task score), every row answered; ahead of GLM-5.3 Flash 73.41, Jev 1.13 71.90, DeepSeek V4.1 Flash 70.68, GPT-5.6 Luna 69.04. The two above are the benchmark team's own Bosun v3.1 1.7B (87.29) and 0.6B (83.20). On the separate Reasoning track: #3, 80.58, behind GLM-5.3 Flash (86.92) and GPT-5.6 Luna (86.17). | leaderboard · registry (results PR #68, merged 28 Sep 2026) |
| Image JevBench v0.1.3 (Benchmark Heaven, released 28 Sep 2026; 49 systems, 684 items) | imajev-4b #1, composite 76.39 (Intelligence 73.8, Calibration 90.5, Speed 87.6, Cost 61.2), ahead of Jev-Omni 73.10 and NeoHorse Jev 4B 71.94; sealed accuracy 83.3% (380/456), second of 44 self-hosted systems; $0.0197 per 1,000 decisions, p50 0.099 s. Measured with this repository at 8501f5c3, adapter c9e5f132, --fast --merge-lora --rotations 1 --calibration calibration.json (it was #11 at 65.72 on v0.1.2 with the slower server). imajev-2b #6 (68.72). The board notes that v0.1.3's 333 fresh sealed items are its own synthetic images. | board · data |
Official JevBench setup: adapter mohit67890/imajev-4b at revision c9e5f132, this repository at a0134749, served with
--rotations 1 --calibration calibration.json (one option order, the shipped calibration file).
The numbers below are from our own runs on 2026-09-24 onward with the released adapters; the raw outputs and per-panel reports are in results/.
JevBench is a text-only benchmark; these are its public splits (111 hard / 72 original / 48 easy) run with the jevbench harness
and the typesafe adapter on one H100, as shipped (four option orders averaged, calibration file applied) unless stated; the pod was busy with other runs, so latencies are under load. They are not
the official board, which adds sealed items and scores four axes (above).
| Panel | imajev-2b | imajev-4b | imajev-9b | Same-protocol references |
|---|---|---|---|---|
| ImajevBench v2.0-lite test (279), 95% cluster CI | 71.7% [0.65, 0.78] | 83.9% [0.79, 0.89] | 82.1% [0.76, 0.88] | untuned bases 60.2 / 70.6 / 76.7%; Jev-Omni 78.5% (its own API); other small VLMs in bench/LEADERBOARD.md |
| · correct Unknown (21) / false abstention (258) | 5 / 4 | 18 / 9 | 15 / 2 | |
| · hidden split (202 items, aggregates only) | 74.3% | 85.6% | 84.7% | |
| JevBench hard (111) | 60.4% | 72.1% | 69.4% | JevK5 v0.2.0 73.9%, Eikos-4B 73.9%, Hopper 67.6%, Qwen3.5-4B base (structured generation) 48.6%, mojev 0.85B 33.3% |
| DecisionBench 1.0 full suite (23,900 rows, the benchmark's own harness, 4 rotations + calibration) | 79.7% (every row scored; #3 of 60 in the official registry; previous version 77.5%) | Bosun v3.1 1.7B 84.9%, 0.6B 81.2%, Winnow-12B 76.7%, Jev 1.13 72.0%; official record merged (Hanno-Labs/decision-bench-results#68), see results/benchmarks/decisionbench/ | ||
| fastino/fast-decisions dev split (1,700 texts, 17 domains, 2,900 classification heads; their board scores a held-out test split) | 60.4% domain macro, 59.4% pooled (previous version 59.3 / 58.8) | not comparable to their board; runner, scorer and predictions in results/benchmarks/fast-decisions/ | ||
| S1-Bench, typed conversion (212 of the 220 English items; our derivative with written distractors, not an S1-Bench score) | 99.1%, ECE 0.019, no abstentions (previous version 98.6%) | saturated: a check that simple questions did not regress; conversion and predictions in results/benchmarks/s1bench-typed/ | ||
| LocalLLaMA/typed-decisions test (400 workflow cases × 5 questions = 2,000 decisions; gold is that dataset's teacher agreement) | 69.2% (calibrated Brier 0.423, ECE 0.025; earlier version 67.0) | Intern-Decision-4B 80.6%, Jev 1.13 73.4%, JevK5 64.5% (Intern-Decision's own runs); ours in results/benchmarks/typed-decisions/ | ||
| Atlan Decision Bench bench-v4 (1,071 rows, 35 tasks from 36 public datasets; their harness, text-only adapter) | 86.6% (928/1,071; 88.6% excluding the 30 icon rows no text-only model can answer); ECE 0.021 | Jev 1.13 92.4%, Claude Haiku 4.5 90.6%, Tev1-4B 85.4%, Laya 52.8% (their runs); 1 training-overlap row disclosed; run in results/benchmarks/atlan-decision-bench/ | ||
| JevBench original (72) / easy (48) | 93.1 / 100 | 98.6 / 100 | 100 / 100 | JevK5 97.2 / 100, Eikos-4B 93.1 / 100, Hopper 95.8 / 100 |
JevBench hard ECE, raw → as shipped (rotations + calibration.json) | 0.176 → 0.123 | 0.113 → 0.082 | 0.187 → 0.092 | JevK5 0.073, Eikos-4B 0.054, Hopper 0.050 |
| MMLU-1000, text-only / with an unrelated photo | 59.8 / 54.9 | 74.5 / 72.9 | 79.2 / 78.8 | previous adapters; not re-run on the shipped versions |
| Irrelevance panel (2,823) | 68.9% | 80.2% | 84.0% | |
| typed-decisions test (2,000) | 59.2% | 67.0% | 67.0% | previous adapters; not re-run on the shipped versions |
| Reasoning dev (6,240; also used for checkpoint selection) | 58.9% | 66.6% | 67.4% | previous adapters; the soft-target checkpoints inside the shipped averages score 62.7 / 67.2 / 68.9% and the averages were not measured; before the last part of the hard-question stage: 64.5 / 67.8 / 69.2% |
Reading:
predict() API. The 4B/9B lead of about 4 points is not significant (p ≈ 0.25; previous adapters vs Jev-Omni) and comes from abstaining on Unknown items,
which Jev-Omni has no output for; on answerable items Jev-Omni is slightly ahead (219 vs 215 / 216 for the previous adapters) and better calibrated
(ECE 0.069). Details in bench/LEADERBOARD.md.unknown code) at that
position are the decision, read through a float32 head. One prefill per request, one forward pass per question, no decoding.--calibration. Temperature never changes an answer, only its probability.unknown is a first-class option in training and inference. Insufficient evidence, a false premise, a mismatched
reference or an answer outside the listed options all train toward unknown.
About a million training decisions in four stages, on open base models, for about $676 of rented GPU time for the whole project. The 2B and 4B went through all four stages (the 4B in one combined run of stages 1 and 2); the 9B, which produced the stage-2 labels, went from stage 1 to stage 3, then stage 4 with the others.
unknown at least 0.5). 71,630 state-grounded (photo vs record) and two-image (reference vs target) decisions
were added (54,468 labelled by the 9B, 17,162 by construction).caiovicentino1/eikos-decisions, CC-BY-4.0, attribution and per-source licences in docs/eikos-decisions-usage.md; only its programmatic and human-annotated rows, nothing labelled by an API model) and 5,000 replayed
image decisions. Soft cross-entropy on the teacher distribution, a rationale loss (weight 0.3, at most 192 tokens) and option
permutation; 2 epochs, lr 2e-5, continuing from the stage-3 adapters (2B 620, 4B 747, 9B 1,018 steps). The checkpoint on its own gained on
JevBench hard but lost photo-plus-record items on ImajevBench, so what ships is the element-wise average of the stage-3 adapter and
this stage's best checkpoint (LoRA and readout): it keeps the image scores and most of the hard-question gain, and passes every
release gate (ImajevBench within 1 point of stage 3, visual and joint item counts, probes, correct-unknown rate, false abstention,
irrelevance).Every teacher is open-weight. No JevBench items (8-gram contamination lint), no Jev outputs and no paid-API outputs were used in training. Total rented GPU across the project: about $676 on RunPod.
The pseudo-labelling pipeline (scripts/v2/pseudo_label.py) and the hard-question pipeline (scripts/p2/) work on your own data.
| imajev-2b | imajev-4b | imajev-9b | |
|---|---|---|---|
| Base (pinned revision) | Qwen/Qwen3.5-2B @15852e8c | Qwen/Qwen3.5-4B @851bf6e8 | Qwen/Qwen3.5-9B @c2022362 |
| Trainable parameters (LoRA + readout) | 16,152,576 | 31,127,040 | 41,152,512 |
Adapter file (adapter_model.safetensors, F32) | 62.6 MB | 122.0 MB | 160.5 MB |
| Shipped adapter | weight-space average (½ + ½) of two adapters: the hard-question adapter and its soft-target continuation | same | same |
| Readout (bias-free linear, float32) | 255 × 2048, 2.1 MB | 255 × 2560, 2.6 MB | 255 × 4096, 4.2 MB |
| Calibration temperature | 1.646 | 1.305 | 1.748 |
| Base weights to download | 4.6 GB | 9.3 GB | 19.3 GB |
Data by stage.
Training. LoRA r16/α32 on all language-model projections including DeltaNet (vision tower frozen), AdamW with weight decay 0, linear warm-up then cosine decay to 10% of the peak, gradient clipping 1.0, 4 GPUs (H100 or H200). Peak learning rates: stage 1 1e-4 for the 2B (after a 2e-4 initial run) and 2e-4 for the 9B; the 4B's first run 1.5e-4; stage 2 5e-5; stage 3 3e-5, then 2e-5; stage 4 2e-5 (soft cross-entropy plus a rationale loss of 0.3, option permutation), followed by the 50/50 weight-space average with the stage-3 adapter.
Compute. About $676 of rented GPU time on RunPod for the whole project, every run included: $499.07 through stage 3 and about $177 for stage 4 (6 h 20 min on one 8×H100 pod).
Full specification: docs/technical-specification.md
JevBench is text-only, so imajev ships a benchmark that is not. ImajevBench v2.0-lite has text-only, visual and joint (photo +
state) items with an explicit Unknown reference; the test split is 279 items in 89 evidence clusters, 21 with an Unknown reference.
It reports cluster-bootstrap CIs and pre-registered paired tests, and ranks direct option scoring and structured generation
separately. It is a preview: all images are AI-generated and there has been no human audit yet. Data and datasheet in bench/,
harness in src/imajev_bench, leaderboard in bench/LEADERBOARD.md. Run your model and send the row.
scripts/v1_text/common.py). New photo sources: PD12M
(CC0), Wikimedia Commons (CC-BY-4.0, CC-BY-3.0 or CC0, checked per file), Open Images (CC-BY-2.0). 16 of the 21 image sources from stage 1 are
admitted under their annotation licences only, with the photos remaining under their upstream terms (not redistributed); for abo,
vizwiz, vizwiz_quality and defects the grant covers the images too. The receipts
and the per-source table are in results/.--calibration and --rotations 4.calibration-modality.json /
calibration-rot4-modality.json add a photo-only bucket (1.028) that the server applies when a request has images and an empty state
(0.012); every other number is unchanged. The 2B and 9B still ship the single temperature. Fit report: results/calibration-modality/.src/vision_decision/ request contracts, MLX backend, Jev API translation, calibration.scripts/playground/ the local server and the playground UI.scripts/train_decision_lora_torch.py, scripts/torch_decision.py training and the PyTorch path.scripts/v2/ data collection, templates and the 9B pseudo-labelling pipeline.scripts/p2/ hard typed-question generation, answering, assembly and temperature fitting.src/imajev_bench/, bench/ the benchmark.results/ every evaluation we report: results/imajev-1.0/ (the released models, their calibration files and release gates; previous-version/ holds the previous adapters and their earlier paired tests), results/benchmarks/
(JevBench and ImajevBench runs), plus reports on earlier checkpoints (results/earlier-checkpoints/).docs/ specs and the run book.imajev is built and maintained by Mohit Garg (mohit67890 on GitHub and Hugging Face),
with Claude (Anthropic) as a co-author on the code. To cite it, see CITATION.cff.
Qwen3.5 (Alibaba) for the base models; Qwen3.6 and gpt-oss (OpenAI) as open-weight teachers; TypeSafe's Jev documentation for the request contract this project mirrors; JevBench (fstandhartinger/jevbench) for the public text splits; kev (jaredpalmer/kev) as the sibling text-only project whose recipe notes were useful; PD12M (Spawning), Wikimedia Commons and Open Images for photos.
Python
90.0%
Shell
4.3%
HTML
2.7%
JavaScript
2.2%
Open Jev-style typed-decision model that also takes images: photo + app state + typed questions in, calibrated probabilities out, locally.
Python
69
14 commits
updated Sep 28, 2026
Small open models (2B · 4B · 9B) that read the photos, records and text a business already has and answer in the options you set,
with a probability on each and an explicit can't tell. Your system acts when it is sure and hands the rest to a person.
Live demo · Website · Quickstart · Checked examples · Results · Technical report
Screenshots of the official leaderboards, captured 28 Sep 2026. Each board is run by its own maintainer, not by us; click through for the live page.
Scored 27 Sep 2026. imajev-4b 67.37, ahead of Plumb-4B (65.84) and TypeSafe's own Jev 1.13.0 (63.29).
Released 28 Sep 2026. imajev-4b 76.39, ahead of Jev-Omni (12B, 73.10) and NeoHorse Jev 4B (71.94); imajev-2b is #6. Sealed accuracy 83.3%, second of 44 self-hosted systems. Measured by the maintainer on the fast server (--fast --merge-lora, one option order, shipped calibration).
"Imajev-4B leads the Jev-class systems with 66.3" (Capability: intelligence and calibration, among systems within 2× Jev's cost and latency).
79.65, ahead of GLM-5.3 Flash (320B), Jev 1.13, DeepSeek V4.1 Flash (552B) and GPT-5.6 Luna. The two above are the benchmark team's own Bosun models.
80.58, behind GLM-5.3 Flash and GPT-5.6 Luna, ahead of DeepSeek V4.1 Flash (552B) and Jev 1.13.
Boards move as new systems are added; ranks are quoted with the board version and the date. Archived copies and the raw data are linked in Official leaderboards.

The live demo runs imajev-4b on a GPU. Pick one of the checked examples (a listing against its photo, a return, a part on the line, an email against a CRM record, a refund against the policy, a ticket, a stylist request it declines to guess on), change the record or swap a photo, and watch the answer and the app's action change. Or upload your own photo and write your own questions.
Send the evidence and the questions you care about. Each answer comes back as a probability over the answers you allowed, ready
for an if. This is the exact script we ran against imajev-4b (site/showcase/listing.py) and its output, with numbers rounded to
three places and usage shortened; 1.15 s on a Mac Studio (four option orders averaged, calibration file applied).
import json, requests
URL = "http://127.0.0.1:8765/v1/systemone"
listing = {
"title": "Men's suede boat shoes",
"color": "red",
"product_type": "shoe",
}
questions = {
"contradicted_field": {
"type": "choice",
"instructions":
"Which field of `listing` does this photo contradict?",
"criteria": {
"listing.color": None,
"listing.product_type": None,
"none of these": "the photo agrees with every field",
},
},
"color_matches": {
"type": "noul",
"instructions":
"The product in the photo matches `listing.color`.",
},
"type_matches": {
"type": "noul",
"instructions": "The photo shows the kind of product "
"given in `listing.product_type`.",
},
}
request = {"state": {"listing": listing}, "questions": questions}
with open("listing.jpg", "rb") as photo:
r = requests.post(URL, files={"image": photo},
data={"request": json.dumps(request)})
print(json.dumps(r.json(), indent=2))
{
"model": "imajev-4b",
"answers": {
"contradicted_field": {
"type": "choice",
"choice": "listing.color",
"probabilities": {
"listing.color": 0.95,
"listing.product_type": 0.006,
"none of these": 0.043
},
"confidence": 0.919,
"unknown_probability": 0.007,
"abstained": false
},
"color_matches": {
"type": "noul",
"noul": 0.082,
"unknown_probability": 0.022,
"abstained": false
},
"type_matches": {
"type": "noul",
"noul": 0.989,
"unknown_probability": 0.004,
"abstained": false
}
},
"usage": {
"total_ms": 1152.9,
"input_tokens": 224
}
}

The same endpoint with no photo: one request routes the ticket (choice), flags urgency (yes/no) and scores frustration (score). Script: site/showcase/ticket.py.
Jev reads text only. General vision models answer in prose, from a hosted API, in seconds. imajev brings Jev's typed answers to photos, and runs on your own hardware.

unknown. Asked for a white or beige bag from a closet with only
a red and a black backpack, the 4B puts 0.89 on unknown instead of guessing.images, unknown_probability and
abstained. Text-only Jev requests work unchanged.| imajev | Jev (TypeSafe) | Jev-Omni | Frontier vision APIs | |
|---|---|---|---|---|
| photos in a request | up to 2 (reference + target) | none, text only | yes; two-photo requests not documented | yes |
| answer format | probability per option you set | probability per option you set | probability per option you set | generated text or JSON |
| says it can't tell | trained unknown, 18 / 21 on ImajevBench (4B) | not documented | no abstain output, so 0 / 21 | only if prompted |
| record per request | 32 KB (about 8k tokens) | 32k tokens | not documented | large |
| where it runs | your hardware, open weights | hosted API | your hardware, open weights (12B) | hosted API |
| time per decision | about 0.1 s raw, about 0.35 s as shipped (one H100) | not documented | about 0.1 s (one H100) | 5 to 8 s |
| ImajevBench accuracy | 83.9% (4B) | cannot take photos | 78.5% | 91.4% to 99.6% |
Jev from docs.typesafe.ai (models page, Jev 1.13.0). Jev-Omni from its model card and our run of its own predict() API.
Frontier rows and all ImajevBench numbers from our runs, 24 Sept 2026; frontier models answer by structured generation, a different
interface. Jev's record limit is larger than imajev's.
The decisions it handles well are the high-volume, well-defined ones with a clear set of answers.
| Area | Decisions (each a checked example on the website) |
|---|---|
| Marketplaces and retail | listing matches its photo · return is the item we shipped · tag a product from a photo |
| Manufacturing and field work | reject a chipped part against a known-good reference · spot what changed on site |
| Customer support | route a ticket and flag urgency · refund against the policy · send a review to the right team |
| Trust and safety | remove spam, harassment and doxxing · ask for a better photo |
| Back office and records | email contradicts the CRM · is the invoice paid? · catalogue record is wrong |
You choose how sure the model must be before it acts. A higher bar automates less and makes fewer mistakes; everything below it goes to a person, including when it says it can't tell.

| Act when at least… | imajev-2b | imajev-4b | imajev-9b |
|---|---|---|---|
| 80% sure | 49% automated, 92.7% right | 63% automated, 94.9% right | 77% automated, 87.9% right |
| 90% sure | 38% automated, 95.3% right | 58% automated, 97.5% right | 70% automated, 91.8% right |
| 99% sure | 21% automated, 100% right | 40% automated, 100% right | 52% automated, 99.3% right |
The 279 ImajevBench test questions (photos, records and text; 21 whose honest answer is can't tell), raw probabilities, scored with the benchmark's own rule (as shipped: four option orders, calibration file). The benchmark is built to be hard, so treat these as a starting point and measure on a few hundred of your own cases before choosing.
Every example comes from one of five small apps in scripts/playground/. Every combination a visitor can click in them is sent to
the model and compared with the right answer (node scripts/playground/verify_scenarios.mjs); a check passes when the answer is
right and the app takes the expected action at an 80% threshold. imajev-4b, served without its calibration file (four option orders):
| App | What it asks | Passed | With the calibration file | |
|---|---|---|---|---|
![]() | Business checks | Does the photo match the listing? Is the return the item we shipped? Is this part chipped? | 19 / 21 | 19 / 21 |
![]() | Text only | An email against a CRM record, a refund against the policy, a post against forum rules, a review, an inbox. Written for the launch and run once. | 20 / 21 | 18 / 21 |
![]() | Wardrobe | Is this what I ordered, does it meet the dress code, do I already own it, which shoes match? | 48 / 49 | 38 / 49 |
![]() | Stylist app | Reads a piece of clothing, then picks bottoms, shoes and a bag from your closet in your colours. | 21 / 24 | 21 / 24 |
![]() | Tracing pad | Reads which letter or number a child traced; the app checks the strokes covered every line. | 22 / 30 | 22 / 30 |
| All | 129 / 145 | 113 / 145 |
With the calibration file the top answer never changes, but confidence is lower, so more cases go to a person at 80%. The misses
worth knowing: with a blank payment note the 4B answered "not paid" instead of unknown, and it reads 8 of 10 scribbles on the tracing
pad as letters. Every run, pass or miss, is in results/scenarios/.
2026-09-26: the 4B moved to its phase-3 adapter (rank-64 LoRA, 256-code readout, trained on the decisions the previous release got wrong): ImajevBench 82.4 → 83.9%, hidden split 84.2 → 85.6%, JevBench hard as shipped 70.3 → 72.1%, DecisionBench full suite 77.5 → 79.7%; it abstains on 18 of 21 ImajevBench Unknown items (was 14) and on 9 of 258 answerable ones (was 3). Its calibration on DecisionBench is worse than before (ECE 0.024 → 0.069). The 2B and 9B are unchanged. Details:
results/phase3/benchmarks.md.
| imajev-2b (latency) | imajev-4b (recommended default) | imajev-9b (quality) | |
|---|---|---|---|
| Base | Qwen3.5-2B (Apache-2.0) | Qwen3.5-4B (Apache-2.0) | Qwen3.5-9B (Apache-2.0) |
| Adapter | LoRA r16/α32 on the language layers + 255-code decision readout; vision tower frozen; shipped as a weight-space average of two adapters (the hard-question adapter and a soft-target continuation of it) | same | same |
| ImajevBench v2.0-lite test | 71.7% | 83.9% | 82.1% |
| JevBench public hard (111), as shipped (4 rotations + calibration) | 60.4% | 72.1% | 69.4% |
| p50 per decision, JevBench hard item, 1×H100, serial, under load | 238 ms as shipped (83 ms raw) | 350 ms (96 ms raw) | 316 ms (96 ms raw) |
| Use it when | latency or memory is the constraint | almost always: within noise of the 9B on ImajevBench | knowledge-heavy text questions, and memory is not a constraint (~19 GB resident) |
| Weights | mohit67890/imajev-2b | mohit67890/imajev-4b | mohit67890/imajev-9b |
Start with the 4B. On ImajevBench it is ahead of the 9B (83.9% vs 82.1%; not significant, paired test p = 0.57) and one point ahead of it on JevBench hard. The 2B is 11 points lower on ImajevBench and 10 lower on JevBench hard; pick it when its footprint is the point. The Mac (MLX) weights agree with the GPU run on 97 to 99% of ImajevBench answers (2B 97.1%, 4B 98.2%, 9B 99.3%).
git clone https://github.com/mohit67890/imajev && cd imajev
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e ".[serve,mlx]" # Apple silicon; elsewhere: pip install -e ".[serve,torch]"
python scripts/download_model.py --model 4b # pinned Qwen3.5-4B into .cache/
hf download mohit67890/imajev-4b --local-dir adapters/imajev-4b
PYTHONPATH=src:scripts python scripts/playground/server.py --model-bundle artifacts/model-qwen4b.json \
--adapter adapters/imajev-4b/mlx --calibration adapters/imajev-4b/calibration.json --model-name imajev-4b --port 8765
Open http://127.0.0.1:8765/ for the playground, or call the API:
curl -s http://127.0.0.1:8765/v1/systemone \
-F 'request={"state":{"listing":{"title":"Blue ceramic mug, 350 ml","colour":"blue"}},
"questions":{"matches":{"type":"noul","instructions":"Does the photo show the listed item?"},
"wrong_field":{"type":"choice","instructions":"Which listing field does the photo contradict?",
"criteria":{"title":null,"colour":null,"none":null}}}}' \
-F image=@photo.jpg
import requests
r = requests.post("http://127.0.0.1:8765/v1/systemone", json={
"state": "Ticket: 'Charged twice for one order, need the duplicate refunded.'",
"questions": {"queue": {"type": "choice", "instructions": "Route the ticket.",
"criteria": {"billing": None, "shipping": None, "account": None, "other": None}},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["can wait a week", "this week", "today", "within the hour", "right now"]}}})
print(r.json()["answers"]["queue"]) # {"type": "choice", "choice": "billing", "probabilities": {...}, "confidence": ..., "unknown_probability": ..., "abstained": false}
On Linux / CUDA, add --backend torch and pass --adapter adapters/imajev-4b (the PEFT adapter at the repo root). For the other
sizes, swap 4b for 2b or 9b in the download commands, the bundle (artifacts/model-qwen9b.json for the 9B; the 2B is the
default bundle) and the adapter paths. --rotations 4 averages four option orders; every number in this README was measured with it and with --calibration (on JevBench hard it adds +1.8 / +0.9 / +0.0 points for the 2B / 4B / 9B at about 3×
the latency). The 9B needs ~19 GB resident; do not keep it and another model loaded on the same Mac.
On a CUDA GPU, --fast makes the torch backend quicker without changing what it computes: one tokenization per question,
image normalisation on the GPU (pixels bit-identical to the processor's), and CUDA graphs of the language model recorded at
load (about a minute; needs a C compiler for the Triton kernels, e.g. build-essential). --merge-lora also folds the adapter
into the weights at load (float32 sum, rounded once to bf16). Checked against a float32 reference on JevBench public and 300
ImajevBench images: --fast --merge-lora is as close to it as the default path, at 11 ms instead of 59 ms per standard text
decision and 91 ms instead of 149 ms per image decision on one H100 (results/serving/fast-path-2026-09-27, scripts/bench_fast_path.py).
What comes back. For each question: choice / score return probabilities over your options (summing to 1, given that the
model answers); noul returns P(yes) with half of the unknown mass added, as in Jev. Every answer also has unknown_probability
(mass on the trained unknown: missing, contradictory or out-of-scope evidence), abstained (unknown was the most likely outcome)
and confidence. Limits per request: up to 2 images, a state up to 32 KB, 1 to 8 questions, 2 to 254 options, 2 to 10 score levels.
Acting on it. Act above a threshold you choose from your own error costs; send abstentions and low-confidence answers to a person.
a = response["answers"]["contradicted_field"]
p = a["probabilities"][a["choice"]] * (1 - a["unknown_probability"])
if a["abstained"] or p < 0.85:
route_to_person(a) # the model cannot tell, or is not sure enough
elif a["choice"] == "none of these":
publish()
else:
hold(field=a["choice"])
Choosing a size. Start with imajev-4b. Use the 2B when latency or memory is tight (it abstains less often than it should on unknown items); use the 9B for knowledge-heavy text questions when ~19 GB of weights is fine.
Checking it on your data. Score a few hundred of your own labelled requests (include "cannot tell" cases) before trusting a
threshold. To fit your own temperature, the evaluators in scripts/ write one JSONL row per question with its logits, and
scripts/v1_text/fit_temperature_calibration.py rows.jsonl calibration.json fits one temperature per question type and option count;
serve it with --calibration. The full guide is section 5 of the technical report.
Measured by each benchmark's maintainer, not by us. Boards move; every rank is quoted with its version and date.
| Board | imajev result | Source |
|---|---|---|
| JevBench v1.4.2.2 (Benchmark Heaven, scored 27 Sep 2026; 91 ranked systems) | imajev-4b #1, JevBench Score 67.37 (Intelligence 52.2, Calibration 80.4, Speed 90.6, Cost 59.7). Plumb-4B 65.84, decider-4b v2 64.13, Jev 1.13.0 63.29. #1 under four of the board's five weightings (#2 under speed-heavy 20:60:20, #3 on Intelligence alone); best Calibration of the top 8. Also #1 on the board's Capability ranking of Jev-class systems (mean of Intelligence and Calibration): 66.3 vs Jev 1.13.0 64.7. Cost on the board's estimate: $0.022 per 1,000 decisions (Jev $0.040). | board · data |
| DecisionBench (eng, v1) (Hanno-Labs, 23 tasks, 23,900 rows; 56 models) | imajev-4b #3, 79.65 (mean task score), every row answered; ahead of GLM-5.3 Flash 73.41, Jev 1.13 71.90, DeepSeek V4.1 Flash 70.68, GPT-5.6 Luna 69.04. The two above are the benchmark team's own Bosun v3.1 1.7B (87.29) and 0.6B (83.20). On the separate Reasoning track: #3, 80.58, behind GLM-5.3 Flash (86.92) and GPT-5.6 Luna (86.17). | leaderboard · registry (results PR #68, merged 28 Sep 2026) |
| Image JevBench v0.1.3 (Benchmark Heaven, released 28 Sep 2026; 49 systems, 684 items) | imajev-4b #1, composite 76.39 (Intelligence 73.8, Calibration 90.5, Speed 87.6, Cost 61.2), ahead of Jev-Omni 73.10 and NeoHorse Jev 4B 71.94; sealed accuracy 83.3% (380/456), second of 44 self-hosted systems; $0.0197 per 1,000 decisions, p50 0.099 s. Measured with this repository at 8501f5c3, adapter c9e5f132, --fast --merge-lora --rotations 1 --calibration calibration.json (it was #11 at 65.72 on v0.1.2 with the slower server). imajev-2b #6 (68.72). The board notes that v0.1.3's 333 fresh sealed items are its own synthetic images. | board · data |
Official JevBench setup: adapter mohit67890/imajev-4b at revision c9e5f132, this repository at a0134749, served with
--rotations 1 --calibration calibration.json (one option order, the shipped calibration file).
The numbers below are from our own runs on 2026-09-24 onward with the released adapters; the raw outputs and per-panel reports are in results/.
JevBench is a text-only benchmark; these are its public splits (111 hard / 72 original / 48 easy) run with the jevbench harness
and the typesafe adapter on one H100, as shipped (four option orders averaged, calibration file applied) unless stated; the pod was busy with other runs, so latencies are under load. They are not
the official board, which adds sealed items and scores four axes (above).
| Panel | imajev-2b | imajev-4b | imajev-9b | Same-protocol references |
|---|---|---|---|---|
| ImajevBench v2.0-lite test (279), 95% cluster CI | 71.7% [0.65, 0.78] | 83.9% [0.79, 0.89] | 82.1% [0.76, 0.88] | untuned bases 60.2 / 70.6 / 76.7%; Jev-Omni 78.5% (its own API); other small VLMs in bench/LEADERBOARD.md |
| · correct Unknown (21) / false abstention (258) | 5 / 4 | 18 / 9 | 15 / 2 | |
| · hidden split (202 items, aggregates only) | 74.3% | 85.6% | 84.7% | |
| JevBench hard (111) | 60.4% | 72.1% | 69.4% | JevK5 v0.2.0 73.9%, Eikos-4B 73.9%, Hopper 67.6%, Qwen3.5-4B base (structured generation) 48.6%, mojev 0.85B 33.3% |
| DecisionBench 1.0 full suite (23,900 rows, the benchmark's own harness, 4 rotations + calibration) | 79.7% (every row scored; #3 of 60 in the official registry; previous version 77.5%) | Bosun v3.1 1.7B 84.9%, 0.6B 81.2%, Winnow-12B 76.7%, Jev 1.13 72.0%; official record merged (Hanno-Labs/decision-bench-results#68), see results/benchmarks/decisionbench/ | ||
| fastino/fast-decisions dev split (1,700 texts, 17 domains, 2,900 classification heads; their board scores a held-out test split) | 60.4% domain macro, 59.4% pooled (previous version 59.3 / 58.8) | not comparable to their board; runner, scorer and predictions in results/benchmarks/fast-decisions/ | ||
| S1-Bench, typed conversion (212 of the 220 English items; our derivative with written distractors, not an S1-Bench score) | 99.1%, ECE 0.019, no abstentions (previous version 98.6%) | saturated: a check that simple questions did not regress; conversion and predictions in results/benchmarks/s1bench-typed/ | ||
| LocalLLaMA/typed-decisions test (400 workflow cases × 5 questions = 2,000 decisions; gold is that dataset's teacher agreement) | 69.2% (calibrated Brier 0.423, ECE 0.025; earlier version 67.0) | Intern-Decision-4B 80.6%, Jev 1.13 73.4%, JevK5 64.5% (Intern-Decision's own runs); ours in results/benchmarks/typed-decisions/ | ||
| Atlan Decision Bench bench-v4 (1,071 rows, 35 tasks from 36 public datasets; their harness, text-only adapter) | 86.6% (928/1,071; 88.6% excluding the 30 icon rows no text-only model can answer); ECE 0.021 | Jev 1.13 92.4%, Claude Haiku 4.5 90.6%, Tev1-4B 85.4%, Laya 52.8% (their runs); 1 training-overlap row disclosed; run in results/benchmarks/atlan-decision-bench/ | ||
| JevBench original (72) / easy (48) | 93.1 / 100 | 98.6 / 100 | 100 / 100 | JevK5 97.2 / 100, Eikos-4B 93.1 / 100, Hopper 95.8 / 100 |
JevBench hard ECE, raw → as shipped (rotations + calibration.json) | 0.176 → 0.123 | 0.113 → 0.082 | 0.187 → 0.092 | JevK5 0.073, Eikos-4B 0.054, Hopper 0.050 |
| MMLU-1000, text-only / with an unrelated photo | 59.8 / 54.9 | 74.5 / 72.9 | 79.2 / 78.8 | previous adapters; not re-run on the shipped versions |
| Irrelevance panel (2,823) | 68.9% | 80.2% | 84.0% | |
| typed-decisions test (2,000) | 59.2% | 67.0% | 67.0% | previous adapters; not re-run on the shipped versions |
| Reasoning dev (6,240; also used for checkpoint selection) | 58.9% | 66.6% | 67.4% | previous adapters; the soft-target checkpoints inside the shipped averages score 62.7 / 67.2 / 68.9% and the averages were not measured; before the last part of the hard-question stage: 64.5 / 67.8 / 69.2% |
Reading:
predict() API. The 4B/9B lead of about 4 points is not significant (p ≈ 0.25; previous adapters vs Jev-Omni) and comes from abstaining on Unknown items,
which Jev-Omni has no output for; on answerable items Jev-Omni is slightly ahead (219 vs 215 / 216 for the previous adapters) and better calibrated
(ECE 0.069). Details in bench/LEADERBOARD.md.unknown code) at that
position are the decision, read through a float32 head. One prefill per request, one forward pass per question, no decoding.--calibration. Temperature never changes an answer, only its probability.unknown is a first-class option in training and inference. Insufficient evidence, a false premise, a mismatched
reference or an answer outside the listed options all train toward unknown.
About a million training decisions in four stages, on open base models, for about $676 of rented GPU time for the whole project. The 2B and 4B went through all four stages (the 4B in one combined run of stages 1 and 2); the 9B, which produced the stage-2 labels, went from stage 1 to stage 3, then stage 4 with the others.
unknown at least 0.5). 71,630 state-grounded (photo vs record) and two-image (reference vs target) decisions
were added (54,468 labelled by the 9B, 17,162 by construction).caiovicentino1/eikos-decisions, CC-BY-4.0, attribution and per-source licences in docs/eikos-decisions-usage.md; only its programmatic and human-annotated rows, nothing labelled by an API model) and 5,000 replayed
image decisions. Soft cross-entropy on the teacher distribution, a rationale loss (weight 0.3, at most 192 tokens) and option
permutation; 2 epochs, lr 2e-5, continuing from the stage-3 adapters (2B 620, 4B 747, 9B 1,018 steps). The checkpoint on its own gained on
JevBench hard but lost photo-plus-record items on ImajevBench, so what ships is the element-wise average of the stage-3 adapter and
this stage's best checkpoint (LoRA and readout): it keeps the image scores and most of the hard-question gain, and passes every
release gate (ImajevBench within 1 point of stage 3, visual and joint item counts, probes, correct-unknown rate, false abstention,
irrelevance).Every teacher is open-weight. No JevBench items (8-gram contamination lint), no Jev outputs and no paid-API outputs were used in training. Total rented GPU across the project: about $676 on RunPod.
The pseudo-labelling pipeline (scripts/v2/pseudo_label.py) and the hard-question pipeline (scripts/p2/) work on your own data.
| imajev-2b | imajev-4b | imajev-9b | |
|---|---|---|---|
| Base (pinned revision) | Qwen/Qwen3.5-2B @15852e8c | Qwen/Qwen3.5-4B @851bf6e8 | Qwen/Qwen3.5-9B @c2022362 |
| Trainable parameters (LoRA + readout) | 16,152,576 | 31,127,040 | 41,152,512 |
Adapter file (adapter_model.safetensors, F32) | 62.6 MB | 122.0 MB | 160.5 MB |
| Shipped adapter | weight-space average (½ + ½) of two adapters: the hard-question adapter and its soft-target continuation | same | same |
| Readout (bias-free linear, float32) | 255 × 2048, 2.1 MB | 255 × 2560, 2.6 MB | 255 × 4096, 4.2 MB |
| Calibration temperature | 1.646 | 1.305 | 1.748 |
| Base weights to download | 4.6 GB | 9.3 GB | 19.3 GB |
Data by stage.
Training. LoRA r16/α32 on all language-model projections including DeltaNet (vision tower frozen), AdamW with weight decay 0, linear warm-up then cosine decay to 10% of the peak, gradient clipping 1.0, 4 GPUs (H100 or H200). Peak learning rates: stage 1 1e-4 for the 2B (after a 2e-4 initial run) and 2e-4 for the 9B; the 4B's first run 1.5e-4; stage 2 5e-5; stage 3 3e-5, then 2e-5; stage 4 2e-5 (soft cross-entropy plus a rationale loss of 0.3, option permutation), followed by the 50/50 weight-space average with the stage-3 adapter.
Compute. About $676 of rented GPU time on RunPod for the whole project, every run included: $499.07 through stage 3 and about $177 for stage 4 (6 h 20 min on one 8×H100 pod).
Full specification: docs/technical-specification.md
JevBench is text-only, so imajev ships a benchmark that is not. ImajevBench v2.0-lite has text-only, visual and joint (photo +
state) items with an explicit Unknown reference; the test split is 279 items in 89 evidence clusters, 21 with an Unknown reference.
It reports cluster-bootstrap CIs and pre-registered paired tests, and ranks direct option scoring and structured generation
separately. It is a preview: all images are AI-generated and there has been no human audit yet. Data and datasheet in bench/,
harness in src/imajev_bench, leaderboard in bench/LEADERBOARD.md. Run your model and send the row.
scripts/v1_text/common.py). New photo sources: PD12M
(CC0), Wikimedia Commons (CC-BY-4.0, CC-BY-3.0 or CC0, checked per file), Open Images (CC-BY-2.0). 16 of the 21 image sources from stage 1 are
admitted under their annotation licences only, with the photos remaining under their upstream terms (not redistributed); for abo,
vizwiz, vizwiz_quality and defects the grant covers the images too. The receipts
and the per-source table are in results/.--calibration and --rotations 4.calibration-modality.json /
calibration-rot4-modality.json add a photo-only bucket (1.028) that the server applies when a request has images and an empty state
(0.012); every other number is unchanged. The 2B and 9B still ship the single temperature. Fit report: results/calibration-modality/.src/vision_decision/ request contracts, MLX backend, Jev API translation, calibration.scripts/playground/ the local server and the playground UI.scripts/train_decision_lora_torch.py, scripts/torch_decision.py training and the PyTorch path.scripts/v2/ data collection, templates and the 9B pseudo-labelling pipeline.scripts/p2/ hard typed-question generation, answering, assembly and temperature fitting.src/imajev_bench/, bench/ the benchmark.results/ every evaluation we report: results/imajev-1.0/ (the released models, their calibration files and release gates; previous-version/ holds the previous adapters and their earlier paired tests), results/benchmarks/
(JevBench and ImajevBench runs), plus reports on earlier checkpoints (results/earlier-checkpoints/).docs/ specs and the run book.imajev is built and maintained by Mohit Garg (mohit67890 on GitHub and Hugging Face),
with Claude (Anthropic) as a co-author on the code. To cite it, see CITATION.cff.
Qwen3.5 (Alibaba) for the base models; Qwen3.6 and gpt-oss (OpenAI) as open-weight teachers; TypeSafe's Jev documentation for the request contract this project mirrors; JevBench (fstandhartinger/jevbench) for the public text splits; kev (jaredpalmer/kev) as the sibling text-only project whose recipe notes were useful; PD12M (Spawning), Wikimedia Commons and Open Images for photos.
Python
90.0%
Shell
4.3%
HTML
2.7%
JavaScript
2.2%