The Jeff models are fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code. You describe a situation and list the options in plain words; Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing: this model takes about 24 ms per decision on an RTX PRO 6000 and 60 ms on an Apple M4 Max (MLX).
Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories don't need to appear in the training data; you describe them, and Jeff picks.
What it is, and what it isn't. These are very small models. They make extremely fast, well-calibrated judgement calls between options, and they slot easily into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't match Jev's, which runs on a much larger model. If zero-shot accuracy isn't good enough for your purposes, a short fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.
Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours, the 2B in about 3.5), all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data.
Independent project. Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe.
| Model | Base model | Base model's panel accuracy (untrained) | Jeff's panel accuracy | Calibration error (ECE) | Decision time (RTX PRO 6000) |
|---|---|---|---|---|---|
| Jeff-Qwen3.5-0.8B | Qwen3.5-0.8B | 45.3% | 79.1% | 0.049 | 22 ms |
| Jeff-Qwen3.5-2B (this model) | Qwen3.5-2B | 46.5% | 83.1% | 0.028 | 24 ms |
| Jeff-Gemma4-E2B | Gemma 4 E2B | 62.5% | 81.6% | 0.031 | 29 ms |
| Jev (published) | — | — | 83.0% | ≈0.06 (average of its per-benchmark figures) | 212 ms per call over the API (Doom harness) |
{
"model": "jeff-latest",
"state": {"voice_transcript": "open the engagment leter", "current_screen": "Deal overview"},
"questions": {
"intent": {
"type": "choice",
"instructions": "Which of these does the user want?",
"criteria": {"1": "Engagement letter", "2": "Inbox", "3": "Deal settings"}
}
}
}
The answer is a probability per option ({"1": 0.94, "2": 0.03, "3": 0.03}), the chosen option and a confidence. Three
question types: choice (pick one of up to 255 options), noul (yes/no, returned as a probability) and score (a
point on a scale you describe). Several independent questions in one request are answered together.
4,599 questions from five public benchmarks, plus JevBench's public hard tier (105 items, scored separately):

| Benchmark | Qwen3.5-0.8B untrained | Jeff-Qwen3.5-0.8B | Qwen3.5-2B untrained | Jeff-Qwen3.5-2B | Gemma 4 E2B untrained | Jeff-Gemma4-E2B | Jev (published) | AutoJev-27B (published) |
|---|---|---|---|---|---|---|---|---|
| Overall (5 benchmarks) | 45.3 | 79.1 | 46.5 | 83.1 | 62.5 | 81.6 | 83.0 | 84.9 |
| BBH | 39.5 | 64.0 | 46.0 | 68.0 | 51.3 | 66.4 | 94.3 | 82.8 |
| Financial PhraseBank | 36.0 | 96.4 | 53.4 | 96.3 | 86.0 | 96.1 | 77.0 | 84.2 |
| JudgeBench | 56.6 | 62.6 | 57.4 | 64.6 | 46.9 | 60.6 | 78.6 | 78.9 |
| RAGTruth | 49.1 | 86.1 | 35.9 | 88.9 | 63.8 | 87.4 | 77.3 | 88.9 |
| WinoGrande | 49.2 | 68.6 | 52.2 | 79.0 | 51.0 | 77.4 | 90.7 | 83.3 |
| JevBench hard (separate) | 36.2 | 47.6 | 45.7 | 53.3 | 41.0 | 48.6 | 73.3 | 70.3 |
Bold: the winner of Jeff against Jev in each row: a Jeff score above Jev's published figure, or Jev's figure where it beats every Jeff model. Bold italic: AutoJev-27B where it is the best of all models in the row (on RAGTruth, tied with Jeff-Qwen3.5-2B); it is shown for reference, since the head-to-head comparison is with Jev. The published Jev and AutoJev figures were measured on a different sample of the same benchmarks. Jeff's overall score comes from classification and grounding (Financial PhraseBank, RAGTruth), where it matches or beats the large models; on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays well below them, as you would expect at this size.
To test zero-shot performance on tasks unlike anything in the benchmarks, we had Jeff play three games. Games aren't the ideal zero-shot test, since a game's state isn't typical unstructured data; but they are a common, and fun, way to test a System 1 model. Each turn, the code describes the situation and the legal moves in words, and the model picks one. The options state what each move leads to (Frogger: "you would be hit by a car and lose a life"; Doom: "the nearest monster is a little to your left"), but never which move is right. Each result is 20 episodes, seed 1234; â–¶ opens a video of the run's first episode.
Jeff-Qwen3.5-0.8B playing, zero-shot (the bold row in the table below):
Doom | Frogger | Pac-Man |
| Model | Doom, kills (monster's direction in words) | Frogger, crossings (consequences) | Pac-Man, pellets of 98 (consequences) | |||
|---|---|---|---|---|---|---|
| Random moves | −0.05 | 0 | 11.2 | |||
| Hand-coded rule bot | 6.55 | â–¶ | 10.25 | â–¶ | 94.1 | â–¶ |
| Qwen3.5-0.8B, untrained | 5.0 | â–¶ | 1.0 | â–¶ | 25.8 | â–¶ |
| Jeff-Qwen3.5-0.8B | 6.55 | â–¶ | 10.3 | â–¶ | 57.0 | â–¶ |
| Qwen3.5-2B, untrained | 0.55 | â–¶ | 0.05 | â–¶ | 72.1 | â–¶ |
| Jeff-Qwen3.5-2B | −0.9 | ▶ | 6.0 | ▶ | 41.2 | ▶ |
| Gemma 4 E2B, untrained | −0.55 | ▶ | 0 | ▶ | 3.2 | ▶ |
| Jeff-Gemma4-E2B | 0.55 | â–¶ | 0.15 | â–¶ | 53.2 | â–¶ |
| Jev (published, Doom) | 6.55, told the aiming rule; −0.60 without it | — | — |
Jeff-0.8B decides in 29–49 ms per move on an M4 Max; Jev's published Doom run took 212 ms per call over its API. The two times were not measured on the same hardware.
Median time per decision over the same 200 benchmark questions (about 200 input tokens each), one question at a time, from raw text to probabilities:
| Model | Parameters | Weights (16-bit) | NVIDIA RTX PRO 6000 | Apple M4 Max (MLX) | CPU (32 threads) |
|---|---|---|---|---|---|
| Jeff-Qwen3.5-0.8B | 0.8B | 1.7 GB | 22 ms | 28 ms | 463 ms |
| Jeff-Qwen3.5-2B | 2B | 4.2 GB | 24 ms | 60 ms | 708 ms |
| Jeff-Gemma4-E2B | 2B effective (4.6B stored) | 9.3 GB | 29 ms | — (MLX runs Qwen only) | 1.0 s |
| AutoJev-27B | 27B | ~54 GB | not published | — | — |
| Jev | not disclosed | API only | 114–212 ms per call in published Doom runs, including the network |
git clone https://github.com/firelex/jeff && cd jeff
uv sync # on Apple silicon: uv sync --extra mac
uv run hf download mstrasser/Jeff-Qwen3.5-2B --local-dir Jeff-Qwen3.5-2B
JEFF_CHECKPOINT=Jeff-Qwen3.5-2B PORT=8765 uv run jeff-serve # NVIDIA or CPU
JEFF_BACKEND=mlx JEFF_CHECKPOINT=Jeff-Qwen3.5-2B PORT=8765 uv run jeff-serve # Apple silicon
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @request.json
{"1": "Engagement letter"}, not long IDs, which cost time and add
nothing.Weights: Apache 2.0 (from Qwen3.5-2B by the Qwen team, Alibaba Cloud). The training data mixes datasets under various licences, listed with their sources in docs/data-sources.md. We release the weights and code, not the training data; some sources are share-alike (CC BY-SA).
The Jeff models are fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code. You describe a situation and list the options in plain words; Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing: this model takes about 24 ms per decision on an RTX PRO 6000 and 60 ms on an Apple M4 Max (MLX).
Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories don't need to appear in the training data; you describe them, and Jeff picks.
What it is, and what it isn't. These are very small models. They make extremely fast, well-calibrated judgement calls between options, and they slot easily into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't match Jev's, which runs on a much larger model. If zero-shot accuracy isn't good enough for your purposes, a short fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.
Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours, the 2B in about 3.5), all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data.
Independent project. Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe.
| Model | Base model | Base model's panel accuracy (untrained) | Jeff's panel accuracy | Calibration error (ECE) | Decision time (RTX PRO 6000) |
|---|---|---|---|---|---|
| Jeff-Qwen3.5-0.8B | Qwen3.5-0.8B | 45.3% | 79.1% | 0.049 | 22 ms |
| Jeff-Qwen3.5-2B (this model) | Qwen3.5-2B | 46.5% | 83.1% | 0.028 | 24 ms |
| Jeff-Gemma4-E2B | Gemma 4 E2B | 62.5% | 81.6% | 0.031 | 29 ms |
| Jev (published) | — | — | 83.0% | ≈0.06 (average of its per-benchmark figures) | 212 ms per call over the API (Doom harness) |
{
"model": "jeff-latest",
"state": {"voice_transcript": "open the engagment leter", "current_screen": "Deal overview"},
"questions": {
"intent": {
"type": "choice",
"instructions": "Which of these does the user want?",
"criteria": {"1": "Engagement letter", "2": "Inbox", "3": "Deal settings"}
}
}
}
The answer is a probability per option ({"1": 0.94, "2": 0.03, "3": 0.03}), the chosen option and a confidence. Three
question types: choice (pick one of up to 255 options), noul (yes/no, returned as a probability) and score (a
point on a scale you describe). Several independent questions in one request are answered together.
4,599 questions from five public benchmarks, plus JevBench's public hard tier (105 items, scored separately):

| Benchmark | Qwen3.5-0.8B untrained | Jeff-Qwen3.5-0.8B | Qwen3.5-2B untrained | Jeff-Qwen3.5-2B | Gemma 4 E2B untrained | Jeff-Gemma4-E2B | Jev (published) | AutoJev-27B (published) |
|---|---|---|---|---|---|---|---|---|
| Overall (5 benchmarks) | 45.3 | 79.1 | 46.5 | 83.1 | 62.5 | 81.6 | 83.0 | 84.9 |
| BBH | 39.5 | 64.0 | 46.0 | 68.0 | 51.3 | 66.4 | 94.3 | 82.8 |
| Financial PhraseBank | 36.0 | 96.4 | 53.4 | 96.3 | 86.0 | 96.1 | 77.0 | 84.2 |
| JudgeBench | 56.6 | 62.6 | 57.4 | 64.6 | 46.9 | 60.6 | 78.6 | 78.9 |
| RAGTruth | 49.1 | 86.1 | 35.9 | 88.9 | 63.8 | 87.4 | 77.3 | 88.9 |
| WinoGrande | 49.2 | 68.6 | 52.2 | 79.0 | 51.0 | 77.4 | 90.7 | 83.3 |
| JevBench hard (separate) | 36.2 | 47.6 | 45.7 | 53.3 | 41.0 | 48.6 | 73.3 | 70.3 |
Bold: the winner of Jeff against Jev in each row: a Jeff score above Jev's published figure, or Jev's figure where it beats every Jeff model. Bold italic: AutoJev-27B where it is the best of all models in the row (on RAGTruth, tied with Jeff-Qwen3.5-2B); it is shown for reference, since the head-to-head comparison is with Jev. The published Jev and AutoJev figures were measured on a different sample of the same benchmarks. Jeff's overall score comes from classification and grounding (Financial PhraseBank, RAGTruth), where it matches or beats the large models; on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays well below them, as you would expect at this size.
To test zero-shot performance on tasks unlike anything in the benchmarks, we had Jeff play three games. Games aren't the ideal zero-shot test, since a game's state isn't typical unstructured data; but they are a common, and fun, way to test a System 1 model. Each turn, the code describes the situation and the legal moves in words, and the model picks one. The options state what each move leads to (Frogger: "you would be hit by a car and lose a life"; Doom: "the nearest monster is a little to your left"), but never which move is right. Each result is 20 episodes, seed 1234; â–¶ opens a video of the run's first episode.
Jeff-Qwen3.5-0.8B playing, zero-shot (the bold row in the table below):
Doom | Frogger | Pac-Man |
| Model | Doom, kills (monster's direction in words) | Frogger, crossings (consequences) | Pac-Man, pellets of 98 (consequences) | |||
|---|---|---|---|---|---|---|
| Random moves | −0.05 | 0 | 11.2 | |||
| Hand-coded rule bot | 6.55 | â–¶ | 10.25 | â–¶ | 94.1 | â–¶ |
| Qwen3.5-0.8B, untrained | 5.0 | â–¶ | 1.0 | â–¶ | 25.8 | â–¶ |
| Jeff-Qwen3.5-0.8B | 6.55 | â–¶ | 10.3 | â–¶ | 57.0 | â–¶ |
| Qwen3.5-2B, untrained | 0.55 | â–¶ | 0.05 | â–¶ | 72.1 | â–¶ |
| Jeff-Qwen3.5-2B | −0.9 | ▶ | 6.0 | ▶ | 41.2 | ▶ |
| Gemma 4 E2B, untrained | −0.55 | ▶ | 0 | ▶ | 3.2 | ▶ |
| Jeff-Gemma4-E2B | 0.55 | â–¶ | 0.15 | â–¶ | 53.2 | â–¶ |
| Jev (published, Doom) | 6.55, told the aiming rule; −0.60 without it | — | — |
Jeff-0.8B decides in 29–49 ms per move on an M4 Max; Jev's published Doom run took 212 ms per call over its API. The two times were not measured on the same hardware.
Median time per decision over the same 200 benchmark questions (about 200 input tokens each), one question at a time, from raw text to probabilities:
| Model | Parameters | Weights (16-bit) | NVIDIA RTX PRO 6000 | Apple M4 Max (MLX) | CPU (32 threads) |
|---|---|---|---|---|---|
| Jeff-Qwen3.5-0.8B | 0.8B | 1.7 GB | 22 ms | 28 ms | 463 ms |
| Jeff-Qwen3.5-2B | 2B | 4.2 GB | 24 ms | 60 ms | 708 ms |
| Jeff-Gemma4-E2B | 2B effective (4.6B stored) | 9.3 GB | 29 ms | — (MLX runs Qwen only) | 1.0 s |
| AutoJev-27B | 27B | ~54 GB | not published | — | — |
| Jev | not disclosed | API only | 114–212 ms per call in published Doom runs, including the network |
git clone https://github.com/firelex/jeff && cd jeff
uv sync # on Apple silicon: uv sync --extra mac
uv run hf download mstrasser/Jeff-Qwen3.5-2B --local-dir Jeff-Qwen3.5-2B
JEFF_CHECKPOINT=Jeff-Qwen3.5-2B PORT=8765 uv run jeff-serve # NVIDIA or CPU
JEFF_BACKEND=mlx JEFF_CHECKPOINT=Jeff-Qwen3.5-2B PORT=8765 uv run jeff-serve # Apple silicon
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @request.json
{"1": "Engagement letter"}, not long IDs, which cost time and add
nothing.Weights: Apache 2.0 (from Qwen3.5-2B by the Qwen team, Alibaba Cloud). The training data mixes datasets under various licences, listed with their sources in docs/data-sources.md. We release the weights and code, not the training data; some sources are share-alike (CC BY-SA).