EldanRing/Winnow-12B

Model

Winnow-12B — GGUF

53

19 commits

3 linked in READMEs

updated Oct 5, 2026

See the code

README

Winnow-12B — GGUF

Choose the target model from the download links in this card or use the Winnow Quickstart. The Hub's automatic model-size/architecture summary currently describes the small Gemma-4-12B-IT-Assistant-BF16.gguf file (862 MB), which is an optional MTP assistant and requires the matching target model. It is not the Winnow target. Both the target and assistant have BF16 files here, so use the explicit target filename or Winnow preset instead of relying on the generic :BF16 snippet. Generic Hub llama.cpp snippets do not provide Winnow's /v1/systemone API.

Local Jev-style decisions, chat, and vision. Q8 tested with 64K context and vision.

Winnow-12B fine-tunes Gemma 4 12B IT for typed decisions. Its llama.cpp-based inference server provides /v1/systemone and ordinary /v1/chat/completions from the same loaded model.

  • Typed decisions: ask noul, choice, and score questions against shared state.
  • Shared computation: prefill the state once, fork question branches, and read answer-token logits without generating answer text.
  • Chat and vision: regular chat, streaming, and image inputs through the same server.
  • Three GGUF model downloads: Q8_0 for the tested 64K vision setup, NVFP4 for a smaller footprint with the supported 8K presets, and BF16 for larger-memory systems. Each contains the merged fine-tune; no conversion or separate LoRA adapter is needed.

Inference code · Quickstart · Full benchmark report · Artifact manifest

GGUF downloads

Choose one model file. All three formats run directly in the Winnow llama.cpp-based server. The optional vision projector is listed separately below.

DownloadPrecision and useFile size
Winnow-12B-Q8_0.ggufQ8_0, 8-bit quantized. Recommended for the tested 64K vision setup.12.67 GB / 11.80 GiB
Winnow-12B-BF16.ggufBF16, 16-bit floating point. Larger-memory systems or CPU/GPU offload.23.83 GB / 22.20 GiB
Winnow-12B-NVFP4.ggufNVFP4 quantized. Smaller-footprint Linux/CUDA 8K text/vision presets; see measured tradeoffs below.8.16 GB / 7.60 GiB

BF16 weights alone exceed 16 GB VRAM; the full-offload measurements below apply to Q8_0. Text-only use needs just the chosen model GGUF. Vision needs that model plus mmproj-Winnow-12B.gguf.

Optional vision projector

Download mmproj-Winnow-12B.gguf — 175 MB / 0.163 GiB.

This is the F16 vision projector, not another model or quantization choice. Use it alongside either BF16 or Q8_0 for image inputs; skip it for text-only use. The same projector works with BF16, Q8_0 and NVFP4.

This repository distributes GGUF model weights only. No safetensors shards or conversion step are required. See Quickstart for exact download and launch commands. Verify downloads.

Decision quality

Winnow BF16 and Q8 versus Jev, Kev and Laya on the frozen public decision benchmarks

Both Winnow variants were evaluated on the same frozen inputs and scoring rules; Jev was served through OpenRouter. BF16 and Q8 were served as GGUF models. Evaluation details are in the benchmark report.

ModelJevBench public subset, 231 itemsKev-v9 clean, 1,046 items
Winnow-12B BF1685.28%81.45%
Winnow-12B Q885.71%81.55%
Jev 1.13, hosted via OpenRouter85.71%87.00%

Winnow Q8 matches Jev on this JevBench public subset: 198 of 231 correct. This is a result on that subset, not a claim of universal parity. JevBench here means public-subset accuracy, not the official composite leaderboard score. The complete report includes competitor comparisons, all measured suites, calibration, hardware and evaluation scope.

The BF16 scores evaluate an earlier GGUF export, not a fresh test of the current download. See the artifact manifest for export provenance.

NVFP4 GGUF: footprint and quality tradeoffs

Winnow-12B-NVFP4.gguf contains 8,163,448,416 bytes. It is a GGUF conversion, distinct from safetensors/HF/vLLM NVFP4 exports; runtime and calibration evidence does not transfer across those backends.

12B Q8 and NVFP4 direct text comparison: 13,529 vs 9,229 MiB peak; 3.43 vs 4.98 decisions per second.

Tested on RTX 5070 Ti 16 GB.

On the matched direct-text workload, Q8 versus NVFP4 used 13,529 versus 9,229 MiB peak device memory and served 3.43 versus 4.98 decisions/s. The profile used four concurrent requests, 4K decision/16K chat context and Q8 KV, with vision and MTP off. These sequential measurements exclude startup and report native decisions/s, not generation tokens/s or sustained service capacity.

Matched 12B direct quality: Q8 versus NVFP4; Jev 85.71 versus 83.55%, Kev 81.45 versus 77.82%, Typed 69.90 versus 70.60%.

Historical matched direct panel. Typed measures teacher-label agreement; deltas are NVFP4 minus Q8.

Historical matched direct panelQ8 GGUFNVFP4 GGUF
Jev public,231 decisions198/231 (85.71%)193/231 (83.55%)
Kev-clean,1,046 decisions852/1,046 (81.45%)814/1,046 (77.82%)
Typed teacher agreement,2,000 decisions /400 groups1,398/2,000 (69.90%)1,412/2,000 (70.60%)

The matched Q8 comparator uses 852/1,046 rather than the separate release campaign’s 853/1,046. The Typed panel is also a separate comparison. Previously observed public panels do not establish unseen-data accuracy; native T1 was used.

64K context with vision

Winnow Q8 64K context and vision measurements

The Q8 model and matching projector use full GPU offload, Q8 KV and exclusive memory scheduling in this measured profile.

MeasurementResult
Configured context capacity65,536 positions
Verified shared prefix with an image65,022 positions, including 1,024 image positions
Observed peak device VRAM15.01 GiB
Four questions at near-full context, cold25.00 s
Same four-question request, cached median of three repeats143.0 ms
Short-prompt generation, median of three 512-token runs55.5 tokens/s
Long vision prompt prefill, 62,435 positions2,893.9 tokens/s
Generation following that long prompt, 512 tokens46.9 tokens/s
Time to first token on that cold long prompt21.75 s

These capacity and timing results are separate workloads, not simultaneous service capacity. Full timing definitions are in the benchmark report.

64K includes prompt formatting, image positions, questions, and generated output. In the tested exclusive profile, chat and decision requests share the weights but take turns using their KV contexts; switching can evict a cached prefix.

Optional reasoning and MTP

Direct native decisions are the default. They read candidate-answer logits without generating an explanation. MTP drafts ordinary chat tokens; optional reasoning is a separate client workflow that can add generated context before native decision scoring. Neither option is enabled implicitly. The Winnow inference server provides the launcher, client, pinned runtime and verified model-specific presets.

Build the inference server

The optional presets require Linux/CUDA. Install the server build prerequisites and a compatible CUDA toolkit first.

git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
python3 scripts/build.py --backend cuda --cuda-arch 120

The 12B BF16 GGUF assistant is 861,519,840 bytes. It is converted from Google's official Gemma 4 12B IT assistant, without training. Use this exact assistant with the Q8/NVFP4 presets; another draft model is not interchangeable. Its license and conversion attribution accompany the download in assistant documentation. Direct serving does not require an assistant.

Supported MTP presets

Tested presetContextVisionObserved peak device memory
12b-nvfp4-vision8k-mtp8,192yes11,773 MiB
12b-q8-text8k-mtp8,192no14,927–15,108 MiB

MTP and vision require additional VRAM. Requirements depend on quantization, context and concurrency. See preset compatibility and settings for supported combinations. Larger-memory configurations have not been verified here.

NVFP4 with vision and MTP:

python3 scripts/winnow.py download --model nv4 --vision on --reasoning off --mtp on
python3 scripts/winnow.py serve --model nv4 --vision on --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.

Q8 text with MTP:

python3 scripts/winnow.py download --model q8 --vision off --reasoning off --mtp on
python3 scripts/winnow.py serve --model q8 --vision off --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.

Q8 direct 64K vision remains available without MTP. BF16 has no validated MTP/adaptive preset.

Experimental reasoning for one text question

The client accepts one named question and a text state. It preserves answer mappings, obtains a native direct decision, and may generate ordinary-template context with this same model before native scoring again. There is no separate thinking mode. Only complete eligible reasoning is scored and blended. After a valid direct decision, failed/incomplete reasoning returns the saved policy-calibrated direct distribution; a failed direct call remains an error. Image states, structured states and multi-question adaptive requests are rejected.

PolicyGate on raw native T1Direct / augmented temperaturesCompleted blend
nvfp4-entropy-v1normalized entropy>0.486191989492708031 /150:50
q8-fixed50-v1maxP<0.81 /150:50

Stop the MTP vision server before launching the separate text-only profile:

python3 scripts/winnow.py download --model nv4 --vision off --reasoning on --mtp on
python3 scripts/winnow.py serve --model nv4 --vision off --reasoning on --mtp on
# For reasoning without MTP, use --mtp off for download, serve and decide.

In a second terminal from the same checkout:

python3 scripts/winnow.py decide --model nv4 --reasoning on --mtp on \
  --input examples/adaptive-decision.json

For Q8, use --model q8 in the download, serve and decide commands; the launcher selects q8-fixed50-v1. The policy uses temperature 0 ordinary-template generation, a 75-second client deadline and natural EOS within the default 8K context. Its 100-word instruction is a soft request, not a hard token cap. Reasoning can materially increase latency.

The frozen Q8 confirmation used 64 groups/96 decisions from two existing sources. Source-equal answer accuracy/rating-consensus agreement changed 62.50%→64.06%, paired 95% change interval [-2.34,+5.47]pp; NLL worsened 1.3226→1.3503 and mean CLI latency rose 198→743ms. Quality is inconclusive. This is not evidence of Boolean transfer, unseen-data accuracy,1 pp preservation or general reasoning ability.

Q8 reasoning confirmation: source-equal accuracy/consensus agreement 62.50 to 64.06%; paired 95% change interval −2.34 to +5.47 pp; mean CLI 198 to 743 ms. Quality inconclusive.

Same-source pilot: 64 groups/96 decisions. The quality measure averages within groups, then equally across sources; it is not pooled accuracy. CLI time includes client startup. The interval and worsened NLL do not establish a general benefit.

NVFP4's retained 900-decision/600-group two-source comparison against an incumbent reasoning route changed source-equal agreement 55.667%→55.250%, paired 95% change interval [-1.083,+0.333]pp; mean resident harness arm latency 751.6→660.0ms. This combines calibration and generated context and is not raw direct production latency or isolated causal reasoning benefit.

A different historical text-only T1/maxP<0.8/equal-blend recipe was also measured. It gained on Jev/Kev-clean and lost on Typed teacher agreement. Mean mandatory direct HTTP inside the workflow versus full router cost was 51→347ms for NVFP4 and 37→234ms for E4B over 3,277 decisions each, including non-routed cases/fallbacks. HTTP excludes router overhead; router time includes it. These are workflow means, not an isolated production comparison or an MTP speedup. Public inputs were previously observed.

Compatibility and scope

Presets provide defaults. Override --context (for example 16k), cache, batch and native branches where supported; choose vision, MTP and reasoning with explicit on/off flags. Run python3 scripts/winnow.py presets for short names and estimated memory. Custom settings do not inherit measured calibration or performance claims. See the setup notes for exact settings and supported combinations. Optional serving is tested on Linux/CUDA; Mac and multi-GPU adaptive use are not validated. Sampled memory peaks are indicative, not a sustained-capacity guarantee. Adaptive image calibration and a universal MTP speedup are not established. Direct serving needs neither assistant weights nor a reasoning policy.

GGUF packaging

Each model GGUF contains its language weights, tokenizer, and chat template. BF16 and Q8_0 are exports of the same merged fine-tune. The NVFP4 GGUF quantizes that merged model without new training. These are alternative model files, not parts to combine. The F16 projector supplies vision support. The remaining root configuration/tokenizer files are reference assets; llama.cpp loads the GGUF directly. No separate LoRA adapter is needed.

Training

Winnow is a LoRA fine-tune, released after merging the learned update into the base model. It is not a full-parameter training run.

SettingValue
Basegoogle/gemma-4-12B-it
Base revision707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7
LoRA rank / alpha32 / 64
LoRA dropout0
Adapted projectionsq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
ExportAdapter merged into BF16 weights, then exported as BF16 GGUF and Q8_0 GGUF

The training dataset and training pipeline remain private and are not released. It contains curated typed-decision examples: synthetic scenarios, teacher-supervised examples, and labeled semantic tasks. Task coverage includes routing, policy and rule application, evidence selection, workflow decisions, ordinal judgments, entailment, paraphrase, and answerability. Contrastive examples vary facts that should change an answer.

Training and validation were split; the public benchmark report documents the separate evaluation scope and any known development exposure. No claim is made about excluding public benchmarks from the base model's pretraining data.

Probabilities and scope

The decision server normalizes logits over the supplied answer options. Its entropy-based confidence summarizes concentration within that distribution; it is not a guaranteed probability of correctness. Reported default decision temperature is 1.0, without a separately fitted calibration map.

Chat and image input are functional in the release runtime. A complete paired general-chat quality comparison against the unfine-tuned base was not completed; this release does not claim identical chat quality. Near-full-context retrieval checks establish capacity and operation, not general 64K reasoning quality. Audio/video capability is not evaluated by this release.

Credits and license

Winnow-12B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 12B IT, released under Apache 2.0. See LICENSE and NOTICE.

The separate inference repository builds on llama.cpp by Georgi Gerganov and contributors and preserves its MIT license. Jev-style refers to the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.

conversational
endpoints_compatible
gemma4
gemma4_unified
gguf
image-text-to-text
local-inference
typed-decisions
vision
winnow

EldanRing/Winnow-12B

Model

Winnow-12B — GGUF

53

19 commits

3 linked in READMEs

updated Oct 5, 2026

See the code

README

Winnow-12B — GGUF

Choose the target model from the download links in this card or use the Winnow Quickstart. The Hub's automatic model-size/architecture summary currently describes the small Gemma-4-12B-IT-Assistant-BF16.gguf file (862 MB), which is an optional MTP assistant and requires the matching target model. It is not the Winnow target. Both the target and assistant have BF16 files here, so use the explicit target filename or Winnow preset instead of relying on the generic :BF16 snippet. Generic Hub llama.cpp snippets do not provide Winnow's /v1/systemone API.

Local Jev-style decisions, chat, and vision. Q8 tested with 64K context and vision.

Winnow-12B fine-tunes Gemma 4 12B IT for typed decisions. Its llama.cpp-based inference server provides /v1/systemone and ordinary /v1/chat/completions from the same loaded model.

  • Typed decisions: ask noul, choice, and score questions against shared state.
  • Shared computation: prefill the state once, fork question branches, and read answer-token logits without generating answer text.
  • Chat and vision: regular chat, streaming, and image inputs through the same server.
  • Three GGUF model downloads: Q8_0 for the tested 64K vision setup, NVFP4 for a smaller footprint with the supported 8K presets, and BF16 for larger-memory systems. Each contains the merged fine-tune; no conversion or separate LoRA adapter is needed.

Inference code · Quickstart · Full benchmark report · Artifact manifest

GGUF downloads

Choose one model file. All three formats run directly in the Winnow llama.cpp-based server. The optional vision projector is listed separately below.

DownloadPrecision and useFile size
Winnow-12B-Q8_0.ggufQ8_0, 8-bit quantized. Recommended for the tested 64K vision setup.12.67 GB / 11.80 GiB
Winnow-12B-BF16.ggufBF16, 16-bit floating point. Larger-memory systems or CPU/GPU offload.23.83 GB / 22.20 GiB
Winnow-12B-NVFP4.ggufNVFP4 quantized. Smaller-footprint Linux/CUDA 8K text/vision presets; see measured tradeoffs below.8.16 GB / 7.60 GiB

BF16 weights alone exceed 16 GB VRAM; the full-offload measurements below apply to Q8_0. Text-only use needs just the chosen model GGUF. Vision needs that model plus mmproj-Winnow-12B.gguf.

Optional vision projector

Download mmproj-Winnow-12B.gguf — 175 MB / 0.163 GiB.

This is the F16 vision projector, not another model or quantization choice. Use it alongside either BF16 or Q8_0 for image inputs; skip it for text-only use. The same projector works with BF16, Q8_0 and NVFP4.

This repository distributes GGUF model weights only. No safetensors shards or conversion step are required. See Quickstart for exact download and launch commands. Verify downloads.

Decision quality

Winnow BF16 and Q8 versus Jev, Kev and Laya on the frozen public decision benchmarks

Both Winnow variants were evaluated on the same frozen inputs and scoring rules; Jev was served through OpenRouter. BF16 and Q8 were served as GGUF models. Evaluation details are in the benchmark report.

ModelJevBench public subset, 231 itemsKev-v9 clean, 1,046 items
Winnow-12B BF1685.28%81.45%
Winnow-12B Q885.71%81.55%
Jev 1.13, hosted via OpenRouter85.71%87.00%

Winnow Q8 matches Jev on this JevBench public subset: 198 of 231 correct. This is a result on that subset, not a claim of universal parity. JevBench here means public-subset accuracy, not the official composite leaderboard score. The complete report includes competitor comparisons, all measured suites, calibration, hardware and evaluation scope.

The BF16 scores evaluate an earlier GGUF export, not a fresh test of the current download. See the artifact manifest for export provenance.

NVFP4 GGUF: footprint and quality tradeoffs

Winnow-12B-NVFP4.gguf contains 8,163,448,416 bytes. It is a GGUF conversion, distinct from safetensors/HF/vLLM NVFP4 exports; runtime and calibration evidence does not transfer across those backends.

12B Q8 and NVFP4 direct text comparison: 13,529 vs 9,229 MiB peak; 3.43 vs 4.98 decisions per second.

Tested on RTX 5070 Ti 16 GB.

On the matched direct-text workload, Q8 versus NVFP4 used 13,529 versus 9,229 MiB peak device memory and served 3.43 versus 4.98 decisions/s. The profile used four concurrent requests, 4K decision/16K chat context and Q8 KV, with vision and MTP off. These sequential measurements exclude startup and report native decisions/s, not generation tokens/s or sustained service capacity.

Matched 12B direct quality: Q8 versus NVFP4; Jev 85.71 versus 83.55%, Kev 81.45 versus 77.82%, Typed 69.90 versus 70.60%.

Historical matched direct panel. Typed measures teacher-label agreement; deltas are NVFP4 minus Q8.

Historical matched direct panelQ8 GGUFNVFP4 GGUF
Jev public,231 decisions198/231 (85.71%)193/231 (83.55%)
Kev-clean,1,046 decisions852/1,046 (81.45%)814/1,046 (77.82%)
Typed teacher agreement,2,000 decisions /400 groups1,398/2,000 (69.90%)1,412/2,000 (70.60%)

The matched Q8 comparator uses 852/1,046 rather than the separate release campaign’s 853/1,046. The Typed panel is also a separate comparison. Previously observed public panels do not establish unseen-data accuracy; native T1 was used.

64K context with vision

Winnow Q8 64K context and vision measurements

The Q8 model and matching projector use full GPU offload, Q8 KV and exclusive memory scheduling in this measured profile.

MeasurementResult
Configured context capacity65,536 positions
Verified shared prefix with an image65,022 positions, including 1,024 image positions
Observed peak device VRAM15.01 GiB
Four questions at near-full context, cold25.00 s
Same four-question request, cached median of three repeats143.0 ms
Short-prompt generation, median of three 512-token runs55.5 tokens/s
Long vision prompt prefill, 62,435 positions2,893.9 tokens/s
Generation following that long prompt, 512 tokens46.9 tokens/s
Time to first token on that cold long prompt21.75 s

These capacity and timing results are separate workloads, not simultaneous service capacity. Full timing definitions are in the benchmark report.

64K includes prompt formatting, image positions, questions, and generated output. In the tested exclusive profile, chat and decision requests share the weights but take turns using their KV contexts; switching can evict a cached prefix.

Optional reasoning and MTP

Direct native decisions are the default. They read candidate-answer logits without generating an explanation. MTP drafts ordinary chat tokens; optional reasoning is a separate client workflow that can add generated context before native decision scoring. Neither option is enabled implicitly. The Winnow inference server provides the launcher, client, pinned runtime and verified model-specific presets.

Build the inference server

The optional presets require Linux/CUDA. Install the server build prerequisites and a compatible CUDA toolkit first.

git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
python3 scripts/build.py --backend cuda --cuda-arch 120

The 12B BF16 GGUF assistant is 861,519,840 bytes. It is converted from Google's official Gemma 4 12B IT assistant, without training. Use this exact assistant with the Q8/NVFP4 presets; another draft model is not interchangeable. Its license and conversion attribution accompany the download in assistant documentation. Direct serving does not require an assistant.

Supported MTP presets

Tested presetContextVisionObserved peak device memory
12b-nvfp4-vision8k-mtp8,192yes11,773 MiB
12b-q8-text8k-mtp8,192no14,927–15,108 MiB

MTP and vision require additional VRAM. Requirements depend on quantization, context and concurrency. See preset compatibility and settings for supported combinations. Larger-memory configurations have not been verified here.

NVFP4 with vision and MTP:

python3 scripts/winnow.py download --model nv4 --vision on --reasoning off --mtp on
python3 scripts/winnow.py serve --model nv4 --vision on --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.

Q8 text with MTP:

python3 scripts/winnow.py download --model q8 --vision off --reasoning off --mtp on
python3 scripts/winnow.py serve --model q8 --vision off --reasoning off --mtp on
# For direct decisions and chat without MTP, use --mtp off in both commands.

Q8 direct 64K vision remains available without MTP. BF16 has no validated MTP/adaptive preset.

Experimental reasoning for one text question

The client accepts one named question and a text state. It preserves answer mappings, obtains a native direct decision, and may generate ordinary-template context with this same model before native scoring again. There is no separate thinking mode. Only complete eligible reasoning is scored and blended. After a valid direct decision, failed/incomplete reasoning returns the saved policy-calibrated direct distribution; a failed direct call remains an error. Image states, structured states and multi-question adaptive requests are rejected.

PolicyGate on raw native T1Direct / augmented temperaturesCompleted blend
nvfp4-entropy-v1normalized entropy>0.486191989492708031 /150:50
q8-fixed50-v1maxP<0.81 /150:50

Stop the MTP vision server before launching the separate text-only profile:

python3 scripts/winnow.py download --model nv4 --vision off --reasoning on --mtp on
python3 scripts/winnow.py serve --model nv4 --vision off --reasoning on --mtp on
# For reasoning without MTP, use --mtp off for download, serve and decide.

In a second terminal from the same checkout:

python3 scripts/winnow.py decide --model nv4 --reasoning on --mtp on \
  --input examples/adaptive-decision.json

For Q8, use --model q8 in the download, serve and decide commands; the launcher selects q8-fixed50-v1. The policy uses temperature 0 ordinary-template generation, a 75-second client deadline and natural EOS within the default 8K context. Its 100-word instruction is a soft request, not a hard token cap. Reasoning can materially increase latency.

The frozen Q8 confirmation used 64 groups/96 decisions from two existing sources. Source-equal answer accuracy/rating-consensus agreement changed 62.50%→64.06%, paired 95% change interval [-2.34,+5.47]pp; NLL worsened 1.3226→1.3503 and mean CLI latency rose 198→743ms. Quality is inconclusive. This is not evidence of Boolean transfer, unseen-data accuracy,1 pp preservation or general reasoning ability.

Q8 reasoning confirmation: source-equal accuracy/consensus agreement 62.50 to 64.06%; paired 95% change interval −2.34 to +5.47 pp; mean CLI 198 to 743 ms. Quality inconclusive.

Same-source pilot: 64 groups/96 decisions. The quality measure averages within groups, then equally across sources; it is not pooled accuracy. CLI time includes client startup. The interval and worsened NLL do not establish a general benefit.

NVFP4's retained 900-decision/600-group two-source comparison against an incumbent reasoning route changed source-equal agreement 55.667%→55.250%, paired 95% change interval [-1.083,+0.333]pp; mean resident harness arm latency 751.6→660.0ms. This combines calibration and generated context and is not raw direct production latency or isolated causal reasoning benefit.

A different historical text-only T1/maxP<0.8/equal-blend recipe was also measured. It gained on Jev/Kev-clean and lost on Typed teacher agreement. Mean mandatory direct HTTP inside the workflow versus full router cost was 51→347ms for NVFP4 and 37→234ms for E4B over 3,277 decisions each, including non-routed cases/fallbacks. HTTP excludes router overhead; router time includes it. These are workflow means, not an isolated production comparison or an MTP speedup. Public inputs were previously observed.

Compatibility and scope

Presets provide defaults. Override --context (for example 16k), cache, batch and native branches where supported; choose vision, MTP and reasoning with explicit on/off flags. Run python3 scripts/winnow.py presets for short names and estimated memory. Custom settings do not inherit measured calibration or performance claims. See the setup notes for exact settings and supported combinations. Optional serving is tested on Linux/CUDA; Mac and multi-GPU adaptive use are not validated. Sampled memory peaks are indicative, not a sustained-capacity guarantee. Adaptive image calibration and a universal MTP speedup are not established. Direct serving needs neither assistant weights nor a reasoning policy.

GGUF packaging

Each model GGUF contains its language weights, tokenizer, and chat template. BF16 and Q8_0 are exports of the same merged fine-tune. The NVFP4 GGUF quantizes that merged model without new training. These are alternative model files, not parts to combine. The F16 projector supplies vision support. The remaining root configuration/tokenizer files are reference assets; llama.cpp loads the GGUF directly. No separate LoRA adapter is needed.

Training

Winnow is a LoRA fine-tune, released after merging the learned update into the base model. It is not a full-parameter training run.

SettingValue
Basegoogle/gemma-4-12B-it
Base revision707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7
LoRA rank / alpha32 / 64
LoRA dropout0
Adapted projectionsq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
ExportAdapter merged into BF16 weights, then exported as BF16 GGUF and Q8_0 GGUF

The training dataset and training pipeline remain private and are not released. It contains curated typed-decision examples: synthetic scenarios, teacher-supervised examples, and labeled semantic tasks. Task coverage includes routing, policy and rule application, evidence selection, workflow decisions, ordinal judgments, entailment, paraphrase, and answerability. Contrastive examples vary facts that should change an answer.

Training and validation were split; the public benchmark report documents the separate evaluation scope and any known development exposure. No claim is made about excluding public benchmarks from the base model's pretraining data.

Probabilities and scope

The decision server normalizes logits over the supplied answer options. Its entropy-based confidence summarizes concentration within that distribution; it is not a guaranteed probability of correctness. Reported default decision temperature is 1.0, without a separately fitted calibration map.

Chat and image input are functional in the release runtime. A complete paired general-chat quality comparison against the unfine-tuned base was not completed; this release does not claim identical chat quality. Near-full-context retrieval checks establish capacity and operation, not general 64K reasoning quality. Audio/video capability is not evaluated by this release.

Credits and license

Winnow-12B is an independent fine-tune by EldanRing of Google DeepMind's Gemma 4 12B IT, released under Apache 2.0. See LICENSE and NOTICE.

The separate inference repository builds on llama.cpp by Georgi Gerganov and contributors and preserves its MIT license. Jev-style refers to the typed-decision interface; Winnow is not affiliated with or endorsed by TypeSafe, Google, or llama.cpp.

conversational
endpoints_compatible
gemma4
gemma4_unified
gguf
image-text-to-text
local-inference
typed-decisions
vision
winnow