Choose the target model from the download links in this card or use the
Winnow Quickstart.
The Hub's automatic model-size/architecture summary currently describes the small
Gemma-4-12B-IT-Assistant-BF16.gguf file (862 MB), which is an optional MTP
assistant and requires the matching target model. It is not the Winnow target.
Both the target and assistant have BF16 files here, so use the explicit target
filename or Winnow preset instead of relying on the generic :BF16 snippet.
Generic Hub llama.cpp snippets do not provide Winnow's /v1/systemone API.
Local Jev-style decisions, chat, and vision. Q8 tested with 64K context and vision.
Winnow-12B fine-tunes Gemma 4 12B IT
for typed decisions. Its llama.cpp-based inference server
provides /v1/systemone and ordinary /v1/chat/completions from the same loaded model.
noul, choice, and score questions against shared state.Inference code · Quickstart · Full benchmark report · Artifact manifest
Choose one model file. All three formats run directly in the Winnow llama.cpp-based server. The optional vision projector is listed separately below.
| Download | Precision and use | File size |
|---|---|---|
| Winnow-12B-Q8_0.gguf | Q8_0, 8-bit quantized. Recommended for the tested 64K vision setup. | 12.67 GB / 11.80 GiB |
| Winnow-12B-BF16.gguf | BF16, 16-bit floating point. Larger-memory systems or CPU/GPU offload. | 23.83 GB / 22.20 GiB |
| Winnow-12B-NVFP4.gguf | NVFP4 quantized. Smaller-footprint Linux/CUDA 8K text/vision presets; see measured tradeoffs below. | 8.16 GB / 7.60 GiB |
BF16 weights alone exceed 16 GB VRAM; the full-offload measurements below
apply to Q8_0. Text-only use needs just the chosen model GGUF. Vision needs
that model plus mmproj-Winnow-12B.gguf.
Download mmproj-Winnow-12B.gguf — 175 MB / 0.163 GiB.
This is the F16 vision projector, not another model or quantization choice. Use it alongside either BF16 or Q8_0 for image inputs; skip it for text-only use. The same projector works with BF16, Q8_0 and NVFP4.
This repository distributes GGUF model weights only. No safetensors shards or conversion step are required. See Quickstart for exact download and launch commands. Verify downloads.

Both Winnow variants were evaluated on the same frozen inputs and scoring rules; Jev was served through OpenRouter. BF16 and Q8 were served as GGUF models. Evaluation details are in the benchmark report.
| Model | JevBench public subset, 231 items | Kev-v9 clean, 1,046 items |
|---|---|---|
| Winnow-12B BF16 | 85.28% | 81.45% |
| Winnow-12B Q8 | 85.71% | 81.55% |
| Jev 1.13, hosted via OpenRouter | 85.71% | 87.00% |
Winnow Q8 matches Jev on this JevBench public subset: 198 of 231 correct. This is a result on that subset, not a claim of universal parity. JevBench here means public-subset accuracy, not the official composite leaderboard score. The complete report includes competitor comparisons, all measured suites, calibration, hardware and evaluation scope.
The BF16 scores evaluate an earlier GGUF export, not a fresh test of the current download. See the artifact manifest for export provenance.
Winnow-12B-NVFP4.gguf contains 8,163,448,416 bytes.
It is a GGUF conversion, distinct from safetensors/HF/vLLM NVFP4 exports; runtime
and calibration evidence does not transfer across those backends.

Tested on RTX 5070 Ti 16 GB.
On the matched direct-text workload, Q8 versus NVFP4 used 13,529 versus 9,229 MiB peak device memory and served 3.43 versus 4.98 decisions/s. The profile used four concurrent requests, 4K decision/16K chat context and Q8 KV, with vision and MTP off. These sequential measurements exclude startup and report native decisions/s, not generation tokens/s or sustained service capacity.

Historical matched direct panel. Typed measures teacher-label agreement; deltas are NVFP4 minus Q8.
| Historical matched direct panel | Q8 GGUF | NVFP4 GGUF |
|---|---|---|
| Jev public,231 decisions | 198/231 (85.71%) | 193/231 (83.55%) |
| Kev-clean,1,046 decisions | 852/1,046 (81.45%) | 814/1,046 (77.82%) |
| Typed teacher agreement,2,000 decisions /400 groups | 1,398/2,000 (69.90%) | 1,412/2,000 (70.60%) |
The matched Q8 comparator uses 852/1,046 rather than the separate release campaign’s 853/1,046. The Typed panel is also a separate comparison. Previously observed public panels do not establish unseen-data accuracy; native T1 was used.

The Q8 model and matching projector use full GPU offload, Q8 KV and exclusive memory scheduling in this measured profile.
| Measurement | Result |
|---|---|
| Configured context capacity | 65,536 positions |
| Verified shared prefix with an image | 65,022 positions, including 1,024 image positions |
| Observed peak device VRAM | 15.01 GiB |
| Four questions at near-full context, cold | 25.00 s |
| Same four-question request, cached median of three repeats | 143.0 ms |
| Short-prompt generation, median of three 512-token runs | 55.5 tokens/s |
| Long vision prompt prefill, 62,435 positions | 2,893.9 tokens/s |
| Generation following that long prompt, 512 tokens | 46.9 tokens/s |
| Time to first token on that cold long prompt | 21.75 s |
These capacity and timing results are separate workloads, not simultaneous service capacity. Full timing definitions are in the benchmark report.
64K includes prompt formatting, image positions, questions, and generated output. In the tested exclusive profile, chat and decision requests share the weights but take turns using their KV contexts; switching can evict a cached prefix.
Direct native decisions are the default. They read candidate-answer logits without generating an explanation. MTP drafts ordinary chat tokens; optional reasoning is a separate client workflow that can add generated context before native decision scoring. Neither option is enabled implicitly. The Winnow inference server provides the launcher, client, pinned runtime and verified model-specific presets.
The optional presets require Linux/CUDA. Install the server build prerequisites and a compatible CUDA toolkit first.
git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
python3 scripts/build.py --backend cuda --cuda-arch 120
The 12B BF16 GGUF assistant is 861,519,840 bytes. It is converted from Google's official Gemma 4 12B IT assistant, without training. Use this exact assistant with the Q8/NVFP4 presets; another draft model is not interchangeable. Its license and conversion attribution accompany the download in assistant documentation. Direct serving does not require an assistant.
| Tested preset | Context | Vision | Observed peak device memory |
|---|---|---|---|
12b-nvfp4-vision8k-mtp | 8,192 | yes | 11,773 MiB |
12b-q8-text8k-mtp | 8,192 | no | 14,927–15,108 MiB |
MTP and vision require additional VRAM. Requirements depend on quantization, context and concurrency. See preset compatibility and settings for supported combinations. Larger-memory configurations have not been verified here.
NVFP4 with vision and MTP:
python3 scripts/winnow.py download --model nv4 --vision on --reasoning off --mtp on
python3 scripts/winnow.py serve --model nv4 --vision on --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.
Q8 text with MTP:
python3 scripts/winnow.py download --model q8 --vision off --reasoning off --mtp on
python3 scripts/winnow.py serve --model q8 --vision off --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.
Q8 direct 64K vision remains available without MTP. BF16 has no validated MTP/adaptive preset.
The client accepts one named question and a text state. It preserves answer mappings, obtains a native direct decision, and may generate ordinary-template context with this same model before native scoring again. There is no separate thinking mode. Only complete eligible reasoning is scored and blended. After a valid direct decision, failed/incomplete reasoning returns the saved policy-calibrated direct distribution; a failed direct call remains an error. Image states, structured states and multi-question adaptive requests are rejected.
| Policy | Gate on raw native T1 | Direct / augmented temperatures | Completed blend |
|---|---|---|---|
nvfp4-entropy-v1 | normalized entropy>0.48619198949270803 | 1 /1 | 50:50 |
q8-fixed50-v1 | maxP<0.8 | 1 /1 | 50:50 |
Stop the MTP vision server before launching the separate text-only profile:
python3 scripts/winnow.py download --model nv4 --vision off --reasoning on --mtp on
python3 scripts/winnow.py serve --model nv4 --vision off --reasoning on --mtp on
# For reasoning without MTP, use --mtp off for download, serve and decide.
In a second terminal from the same checkout:
python3 scripts/winnow.py decide --model nv4 --reasoning on --mtp on \
--input examples/adaptive-decision.json
For Q8, use --model q8 in the download, serve and decide commands; the launcher selects q8-fixed50-v1.
The policy uses temperature 0 ordinary-template generation, a 75-second client
deadline and natural EOS within the default 8K context. Its 100-word instruction is a soft
request, not a hard token cap. Reasoning can materially increase latency.
The frozen Q8 confirmation used 64 groups/96 decisions from two existing sources. Source-equal answer accuracy/rating-consensus agreement changed 62.50%→64.06%, paired 95% change interval [-2.34,+5.47]pp; NLL worsened 1.3226→1.3503 and mean CLI latency rose 198→743ms. Quality is inconclusive. This is not evidence of Boolean transfer, unseen-data accuracy,1 pp preservation or general reasoning ability.

Same-source pilot: 64 groups/96 decisions. The quality measure averages within groups, then equally across sources; it is not pooled accuracy. CLI time includes client startup. The interval and worsened NLL do not establish a general benefit.
NVFP4's retained 900-decision/600-group two-source comparison against an incumbent reasoning route changed source-equal agreement 55.667%→55.250%, paired 95% change interval [-1.083,+0.333]pp; mean resident harness arm latency 751.6→660.0ms. This combines calibration and generated context and is not raw direct production latency or isolated causal reasoning benefit.
A different historical text-only T1/maxP<0.8/equal-blend recipe was also measured. It gained on Jev/Kev-clean and lost on Typed teacher agreement. Mean mandatory direct HTTP inside the workflow versus full router cost was 51→347ms for NVFP4 and 37→234ms for E4B over 3,277 decisions each, including non-routed cases/fallbacks. HTTP excludes router overhead; router time includes it. These are workflow means, not an isolated production comparison or an MTP speedup. Public inputs were previously observed.
Presets provide defaults. Override --context (for example 16k), cache, batch
and native branches where supported; choose vision, MTP and reasoning with
explicit on/off flags. Run python3 scripts/winnow.py presets for short names
and estimated memory. Custom settings do not inherit measured calibration or
performance claims. See the setup notes
for exact settings and supported combinations. Optional serving is tested on
Linux/CUDA; Mac and multi-GPU adaptive use are not validated. Sampled memory peaks
are indicative, not a sustained-capacity guarantee. Adaptive image calibration
and a universal MTP speedup are not established. Direct serving needs neither
assistant weights nor a reasoning policy.
Each model GGUF contains its language weights, tokenizer, and chat template. BF16 and Q8_0 are exports of the same merged fine-tune. The NVFP4 GGUF quantizes that merged model without new training. These are alternative model files, not parts to combine. The F16 projector supplies vision support. The remaining root configuration/tokenizer files are reference assets; llama.cpp loads the GGUF directly. No separate LoRA adapter is needed.
Winnow is a LoRA fine-tune, released after merging the learned update into the base model. It is not a full-parameter training run.
| Setting | Value |
|---|---|
| Base | google/gemma-4-12B-it |
| Base revision | 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 |
| LoRA rank / alpha | 32 / 64 |
| LoRA dropout | 0 |
| Adapted projections | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Export | Adapter merged into BF16 weights, then exported as BF16 GGUF and Q8_0 GGUF |
The training dataset and training pipeline remain private and are not released. It contains curated typed-decision examples: synthetic scenarios, teacher-supervised examples, and labeled semantic tasks. Task coverage includes routing, policy and rule application, evidence selection, workflow decisions, ordinal judgments, entailment, paraphrase, and answerability. Contrastive examples vary facts that should change an answer.
Training and validation were split; the public benchmark report documents the separate evaluation scope and any known development exposure. No claim is made about excluding public benchmarks from the base model's pretraining data.
The decision server normalizes logits over the supplied answer options. Its entropy-based confidence summarizes concentration within that distribution; it is not a guaranteed probability of correctness. Reported default decision temperature is 1.0, without a separately fitted calibration map.
Chat and image input are functional in the release runtime. A complete paired general-chat quality comparison against the unfine-tuned base was not completed; this release does not claim identical chat quality. Near-full-context retrieval checks establish capacity and operation, not general 64K reasoning quality. Audio/video capability is not evaluated by this release.
Winnow-12B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 12B IT, released under Apache 2.0. See LICENSE and NOTICE.
The separate inference repository builds on llama.cpp by Georgi Gerganov and contributors and preserves its MIT license. Jev-style refers to the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.
Choose the target model from the download links in this card or use the
Winnow Quickstart.
The Hub's automatic model-size/architecture summary currently describes the small
Gemma-4-12B-IT-Assistant-BF16.gguf file (862 MB), which is an optional MTP
assistant and requires the matching target model. It is not the Winnow target.
Both the target and assistant have BF16 files here, so use the explicit target
filename or Winnow preset instead of relying on the generic :BF16 snippet.
Generic Hub llama.cpp snippets do not provide Winnow's /v1/systemone API.
Local Jev-style decisions, chat, and vision. Q8 tested with 64K context and vision.
Winnow-12B fine-tunes Gemma 4 12B IT
for typed decisions. Its llama.cpp-based inference server
provides /v1/systemone and ordinary /v1/chat/completions from the same loaded model.
noul, choice, and score questions against shared state.Inference code · Quickstart · Full benchmark report · Artifact manifest
Choose one model file. All three formats run directly in the Winnow llama.cpp-based server. The optional vision projector is listed separately below.
| Download | Precision and use | File size |
|---|---|---|
| Winnow-12B-Q8_0.gguf | Q8_0, 8-bit quantized. Recommended for the tested 64K vision setup. | 12.67 GB / 11.80 GiB |
| Winnow-12B-BF16.gguf | BF16, 16-bit floating point. Larger-memory systems or CPU/GPU offload. | 23.83 GB / 22.20 GiB |
| Winnow-12B-NVFP4.gguf | NVFP4 quantized. Smaller-footprint Linux/CUDA 8K text/vision presets; see measured tradeoffs below. | 8.16 GB / 7.60 GiB |
BF16 weights alone exceed 16 GB VRAM; the full-offload measurements below
apply to Q8_0. Text-only use needs just the chosen model GGUF. Vision needs
that model plus mmproj-Winnow-12B.gguf.
Download mmproj-Winnow-12B.gguf — 175 MB / 0.163 GiB.
This is the F16 vision projector, not another model or quantization choice. Use it alongside either BF16 or Q8_0 for image inputs; skip it for text-only use. The same projector works with BF16, Q8_0 and NVFP4.
This repository distributes GGUF model weights only. No safetensors shards or conversion step are required. See Quickstart for exact download and launch commands. Verify downloads.

Both Winnow variants were evaluated on the same frozen inputs and scoring rules; Jev was served through OpenRouter. BF16 and Q8 were served as GGUF models. Evaluation details are in the benchmark report.
| Model | JevBench public subset, 231 items | Kev-v9 clean, 1,046 items |
|---|---|---|
| Winnow-12B BF16 | 85.28% | 81.45% |
| Winnow-12B Q8 | 85.71% | 81.55% |
| Jev 1.13, hosted via OpenRouter | 85.71% | 87.00% |
Winnow Q8 matches Jev on this JevBench public subset: 198 of 231 correct. This is a result on that subset, not a claim of universal parity. JevBench here means public-subset accuracy, not the official composite leaderboard score. The complete report includes competitor comparisons, all measured suites, calibration, hardware and evaluation scope.
The BF16 scores evaluate an earlier GGUF export, not a fresh test of the current download. See the artifact manifest for export provenance.
Winnow-12B-NVFP4.gguf contains 8,163,448,416 bytes.
It is a GGUF conversion, distinct from safetensors/HF/vLLM NVFP4 exports; runtime
and calibration evidence does not transfer across those backends.

Tested on RTX 5070 Ti 16 GB.
On the matched direct-text workload, Q8 versus NVFP4 used 13,529 versus 9,229 MiB peak device memory and served 3.43 versus 4.98 decisions/s. The profile used four concurrent requests, 4K decision/16K chat context and Q8 KV, with vision and MTP off. These sequential measurements exclude startup and report native decisions/s, not generation tokens/s or sustained service capacity.

Historical matched direct panel. Typed measures teacher-label agreement; deltas are NVFP4 minus Q8.
| Historical matched direct panel | Q8 GGUF | NVFP4 GGUF |
|---|---|---|
| Jev public,231 decisions | 198/231 (85.71%) | 193/231 (83.55%) |
| Kev-clean,1,046 decisions | 852/1,046 (81.45%) | 814/1,046 (77.82%) |
| Typed teacher agreement,2,000 decisions /400 groups | 1,398/2,000 (69.90%) | 1,412/2,000 (70.60%) |
The matched Q8 comparator uses 852/1,046 rather than the separate release campaign’s 853/1,046. The Typed panel is also a separate comparison. Previously observed public panels do not establish unseen-data accuracy; native T1 was used.

The Q8 model and matching projector use full GPU offload, Q8 KV and exclusive memory scheduling in this measured profile.
| Measurement | Result |
|---|---|
| Configured context capacity | 65,536 positions |
| Verified shared prefix with an image | 65,022 positions, including 1,024 image positions |
| Observed peak device VRAM | 15.01 GiB |
| Four questions at near-full context, cold | 25.00 s |
| Same four-question request, cached median of three repeats | 143.0 ms |
| Short-prompt generation, median of three 512-token runs | 55.5 tokens/s |
| Long vision prompt prefill, 62,435 positions | 2,893.9 tokens/s |
| Generation following that long prompt, 512 tokens | 46.9 tokens/s |
| Time to first token on that cold long prompt | 21.75 s |
These capacity and timing results are separate workloads, not simultaneous service capacity. Full timing definitions are in the benchmark report.
64K includes prompt formatting, image positions, questions, and generated output. In the tested exclusive profile, chat and decision requests share the weights but take turns using their KV contexts; switching can evict a cached prefix.
Direct native decisions are the default. They read candidate-answer logits without generating an explanation. MTP drafts ordinary chat tokens; optional reasoning is a separate client workflow that can add generated context before native decision scoring. Neither option is enabled implicitly. The Winnow inference server provides the launcher, client, pinned runtime and verified model-specific presets.
The optional presets require Linux/CUDA. Install the server build prerequisites and a compatible CUDA toolkit first.
git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
python3 scripts/build.py --backend cuda --cuda-arch 120
The 12B BF16 GGUF assistant is 861,519,840 bytes. It is converted from Google's official Gemma 4 12B IT assistant, without training. Use this exact assistant with the Q8/NVFP4 presets; another draft model is not interchangeable. Its license and conversion attribution accompany the download in assistant documentation. Direct serving does not require an assistant.
| Tested preset | Context | Vision | Observed peak device memory |
|---|---|---|---|
12b-nvfp4-vision8k-mtp | 8,192 | yes | 11,773 MiB |
12b-q8-text8k-mtp | 8,192 | no | 14,927–15,108 MiB |
MTP and vision require additional VRAM. Requirements depend on quantization, context and concurrency. See preset compatibility and settings for supported combinations. Larger-memory configurations have not been verified here.
NVFP4 with vision and MTP:
python3 scripts/winnow.py download --model nv4 --vision on --reasoning off --mtp on
python3 scripts/winnow.py serve --model nv4 --vision on --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.
Q8 text with MTP:
python3 scripts/winnow.py download --model q8 --vision off --reasoning off --mtp on
python3 scripts/winnow.py serve --model q8 --vision off --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.
Q8 direct 64K vision remains available without MTP. BF16 has no validated MTP/adaptive preset.
The client accepts one named question and a text state. It preserves answer mappings, obtains a native direct decision, and may generate ordinary-template context with this same model before native scoring again. There is no separate thinking mode. Only complete eligible reasoning is scored and blended. After a valid direct decision, failed/incomplete reasoning returns the saved policy-calibrated direct distribution; a failed direct call remains an error. Image states, structured states and multi-question adaptive requests are rejected.
| Policy | Gate on raw native T1 | Direct / augmented temperatures | Completed blend |
|---|---|---|---|
nvfp4-entropy-v1 | normalized entropy>0.48619198949270803 | 1 /1 | 50:50 |
q8-fixed50-v1 | maxP<0.8 | 1 /1 | 50:50 |
Stop the MTP vision server before launching the separate text-only profile:
python3 scripts/winnow.py download --model nv4 --vision off --reasoning on --mtp on
python3 scripts/winnow.py serve --model nv4 --vision off --reasoning on --mtp on
# For reasoning without MTP, use --mtp off for download, serve and decide.
In a second terminal from the same checkout:
python3 scripts/winnow.py decide --model nv4 --reasoning on --mtp on \
--input examples/adaptive-decision.json
For Q8, use --model q8 in the download, serve and decide commands; the launcher selects q8-fixed50-v1.
The policy uses temperature 0 ordinary-template generation, a 75-second client
deadline and natural EOS within the default 8K context. Its 100-word instruction is a soft
request, not a hard token cap. Reasoning can materially increase latency.
The frozen Q8 confirmation used 64 groups/96 decisions from two existing sources. Source-equal answer accuracy/rating-consensus agreement changed 62.50%→64.06%, paired 95% change interval [-2.34,+5.47]pp; NLL worsened 1.3226→1.3503 and mean CLI latency rose 198→743ms. Quality is inconclusive. This is not evidence of Boolean transfer, unseen-data accuracy,1 pp preservation or general reasoning ability.

Same-source pilot: 64 groups/96 decisions. The quality measure averages within groups, then equally across sources; it is not pooled accuracy. CLI time includes client startup. The interval and worsened NLL do not establish a general benefit.
NVFP4's retained 900-decision/600-group two-source comparison against an incumbent reasoning route changed source-equal agreement 55.667%→55.250%, paired 95% change interval [-1.083,+0.333]pp; mean resident harness arm latency 751.6→660.0ms. This combines calibration and generated context and is not raw direct production latency or isolated causal reasoning benefit.
A different historical text-only T1/maxP<0.8/equal-blend recipe was also measured. It gained on Jev/Kev-clean and lost on Typed teacher agreement. Mean mandatory direct HTTP inside the workflow versus full router cost was 51→347ms for NVFP4 and 37→234ms for E4B over 3,277 decisions each, including non-routed cases/fallbacks. HTTP excludes router overhead; router time includes it. These are workflow means, not an isolated production comparison or an MTP speedup. Public inputs were previously observed.
Presets provide defaults. Override --context (for example 16k), cache, batch
and native branches where supported; choose vision, MTP and reasoning with
explicit on/off flags. Run python3 scripts/winnow.py presets for short names
and estimated memory. Custom settings do not inherit measured calibration or
performance claims. See the setup notes
for exact settings and supported combinations. Optional serving is tested on
Linux/CUDA; Mac and multi-GPU adaptive use are not validated. Sampled memory peaks
are indicative, not a sustained-capacity guarantee. Adaptive image calibration
and a universal MTP speedup are not established. Direct serving needs neither
assistant weights nor a reasoning policy.
Each model GGUF contains its language weights, tokenizer, and chat template. BF16 and Q8_0 are exports of the same merged fine-tune. The NVFP4 GGUF quantizes that merged model without new training. These are alternative model files, not parts to combine. The F16 projector supplies vision support. The remaining root configuration/tokenizer files are reference assets; llama.cpp loads the GGUF directly. No separate LoRA adapter is needed.
Winnow is a LoRA fine-tune, released after merging the learned update into the base model. It is not a full-parameter training run.
| Setting | Value |
|---|---|
| Base | google/gemma-4-12B-it |
| Base revision | 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 |
| LoRA rank / alpha | 32 / 64 |
| LoRA dropout | 0 |
| Adapted projections | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Export | Adapter merged into BF16 weights, then exported as BF16 GGUF and Q8_0 GGUF |
The training dataset and training pipeline remain private and are not released. It contains curated typed-decision examples: synthetic scenarios, teacher-supervised examples, and labeled semantic tasks. Task coverage includes routing, policy and rule application, evidence selection, workflow decisions, ordinal judgments, entailment, paraphrase, and answerability. Contrastive examples vary facts that should change an answer.
Training and validation were split; the public benchmark report documents the separate evaluation scope and any known development exposure. No claim is made about excluding public benchmarks from the base model's pretraining data.
The decision server normalizes logits over the supplied answer options. Its entropy-based confidence summarizes concentration within that distribution; it is not a guaranteed probability of correctness. Reported default decision temperature is 1.0, without a separately fitted calibration map.
Chat and image input are functional in the release runtime. A complete paired general-chat quality comparison against the unfine-tuned base was not completed; this release does not claim identical chat quality. Near-full-context retrieval checks establish capacity and operation, not general 64K reasoning quality. Audio/video capability is not evaluated by this release.
Winnow-12B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 12B IT, released under Apache 2.0. See LICENSE and NOTICE.
The separate inference repository builds on llama.cpp by Georgi Gerganov and contributors and preserves its MIT license. Jev-style refers to the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.