Local typed decisions, chat and vision with Winnow-12B and llama.cpp
See the codeA native llama.cpp server for local typed decisions and regular chat, sharing one
loaded model. /v1/systemone evaluates noul, choice, and score questions
against a shared state. /v1/chat/completions retains llama-server's chat, vision,
and streaming interfaces.
Native NVIDIA CUDA and Apple Silicon Metal + Accelerate backends expose the same decision, chat and vision APIs.
This release also contains an explicitly experimental adaptive decision client. Direct decisions remain the default. The optional Linux/CUDA text-only mode uses verified model artifacts and optional matching MTP assistants and frozen model-specific policies; it is not an accuracy-preservation guarantee. See the adaptive contract and limits.
E4B adaptive policy update: the opt-in E4B client now uses the tested
e4b-calibrated75-g95-v1 policy. On 288 text decisions it matched 198 labels
versus 185 with the previous E4B policy, at higher reasoning cost (resident
mean 376 versus 263 ms). The previous E4B policy remains available by ID;
12B Q8 and NVFP4 policies, model weights, and the server binary are unchanged.
Direct decisions remain the default. See results and limits
and upgrade instructions.
64K context and vision on a 16 GB RTX 5070 Ti, with Q8 weights fully on the GPU. Winnow-12B is a fine-tune of Gemma 4 12B; the same loaded model serves typed decisions and regular chat.
Model weights and model card · Benchmarks · API · Installation
On the 231-item public JevBench subset, Winnow-12B Q8 matched hosted Jev at 85.7% (198/231); Winnow-12B BF16 scored 85.3%. Local-model comparisons use the same RTX PRO 5000 Blackwell. See the full benchmark report.

BF16 GGUF, Q8 GGUF, and the matching vision projector are separate downloads. The training dataset is private. Exact release checksums are in the model manifest.
| 12B Q8 method | Correct / 231 | Accuracy |
|---|---|---|
| Native direct | 198 | 85.71% |
| Same-model reasoning, MTP off | 208 | 90.04% |
| Same-model reasoning, MTP draft depth 1 | 209 | 90.48% |
This all-cases, 192-token recipe scored partial text and used no gate or blend;
it is not the released q8-fixed50-v1 policy or its separate 96-decision pilot.
The MTP variant was faster but regressed from 83/96 to 81/96 on an external panel.
Full Q8 benchmark, completion handling, probability losses and timing.
A historical text-only experiment added same-model reasoning below a raw confidence threshold, then blended direct and augmented probabilities equally. It improved Jev and Kev-clean results but reduced Typed teacher agreement (direct → adaptive):
| Historical model / recipe | Jev, 231 | Kev-clean, 1,046 | Typed agreement, 2,000 |
|---|---|---|---|
| 12B NVFP4, fixed gate | 83.55 → 87.88% | 77.82 → 81.45% | 70.60 → 70.15% |
| E4B Q8, fixed gate | 80.52 → 83.55% | 72.66 → 76.96% | 72.35 → 69.50% |
These historical recipes differ from the released NVFP4 entropy and E4B calibrated policies. Full-router means were 347 ms and 234 ms; their included direct HTTP calls averaged 51 ms and 37 ms. Later E4B calibration improved probability losses with little agreement change; the small Q8 confirmation remained inconclusive. See full counts, policy comparisons, latency definitions and MTP limits. Direct decisions remain the default.
Choose direct decisions, reasoning and MTP with one model selector and explicit on/off switches. The guide covers verified downloads, source setup, the thin runtime, and supported combinations for 12B Q8, 12B NVFP4 and E4B Q8.
Supported: Linux + NVIDIA GPU (the measured profile uses a 16 GB RTX 5070 Ti) or Apple Silicon Mac (24 GB or more unified memory for the documented profile). Use a terminal to start the local API server.
git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
Apple Silicon Mac — native Metal + Accelerate:
xcode-select --install
brew install python cmake openssl@3
Wait for Xcode command-line tools to finish installing. The brew command assumes
Homebrew is installed. Setup locates its OpenSSL automatically.
Linux/NVIDIA — native CUDA:
sudo apt-get install build-essential cmake git python3 libssl-dev
Install a compatible NVIDIA driver and the CUDA toolkit. A driver alone is not enough; Blackwell needs CUDA 12.8 or newer. The setup script checks that your toolkit supports your GPU before it downloads anything.
python3 scripts/winnow.py presets
python3 scripts/winnow.py setup --model q8
Short presets are q8, nv4 and e4b. Each defaults to 8K text with MTP and
reasoning off. Setup checks prerequisites, downloads verified weights and builds.
It does not install system packages. Existing valid downloads are reused.
python3 scripts/winnow.py serve --model q8 --context 8k
# In a second terminal:
python3 scripts/winnow.py decide --model q8 --input examples/decisions.json
Select --vision on, --mtp on or --reasoning on explicitly; download the
matching assets first with the same flags. Context and numerical settings can be
overridden where supported. The quickstart includes the mode
matrix, memory estimates, measured baselines and examples. Changed configurations
do not inherit the measured calibration or performance claims.
The server uses http://127.0.0.1:8091. Leave its terminal running; Ctrl+C stops
it. Direct clients verify the chosen model and quantization. Use --model-dir
for another verified asset directory. The existing scripts/serve.py remains
available for custom GGUF paths and legacy platform profiles.
Full installation and troubleshooting ·
API, vision and concurrency · python3 scripts/winnow.py --help
For the default 12B Q8 text-only, 8K setup above, this short release check starts and stops its own authenticated local servers, including selected/full-head comparison. These model paths are Q8-specific; they do not select E4B or NVFP4:
python3 scripts/release_check.py --model models/gguf/Winnow-12B-Q8_0.gguf \
--context 8192 --memory exclusive --output results/release-smoke-text
It uses a 4,096-token output allowance and does not run a quality benchmark or context sweep. For an already running text server, use the individual checks:
python3 scripts/check.py --output results/check
python3 scripts/bench.py --output results/timings --repeats 5
python3 scripts/parity.py --output results/selected
For the optional 12B Q8 64K vision check, first download the projector explicitly. Use a host that meets the 64K platform requirements:
python3 scripts/winnow.py download --model q8 --vision on
python3 scripts/release_check.py --model models/gguf/Winnow-12B-Q8_0.gguf \
--mmproj models/gguf/mmproj-F16.gguf --context 65536 \
--memory exclusive --output results/release-smoke-vision
Add --long-context to that release check explicitly for capacity confirmation.
The following probes require an already running 64K vision server, launched
with the matching platform command in INSTALL:
python3 scripts/check.py --output results/check-vision --vision
python3 scripts/check.py --output results/64k --vision --long-context 65536 --image-size 1536
Authenticated servers are supported through WINNOW_API_KEY_FILE or WINNOW_API_KEY.
Restart with --head full, then use scripts/parity.py --output results/full --compare results/selected to verify selected-head probabilities. The full-head
mode is a reference implementation; normal chat always retains its full head.
All scripts save requests and responses locally. These focused probes validate
runtime behavior and timings, not general model quality or calibrated confidence.
Build the container locally on Linux. GPU execution requires Docker Engine with the NVIDIA Container Toolkit and a CUDA 13.0-compatible driver. Apple Silicon uses native Metal above.
For RTX 50-series / SM120, download both 12B Q8 and its vision projector before starting the default 64K vision container:
python3 scripts/winnow.py download --model q8 --vision on
docker build --build-arg CUDA_ARCH=120 -t winnow-inference .
docker run --rm --gpus 'device=0' -p 127.0.0.1:8091:8091 \
-v "$PWD/models:/models:ro" winnow-inference
The default is the 5070 Ti 64K + vision profile. For a complete container-only
setup, including downloading weights without host Python, follow
the container quickstart.
Other NVIDIA architectures can build the Dockerfile with their own CUDA_ARCH.
The runtime image is about 1.7 GB uncompressed, excludes model weights and build tools, and runs as an unprivileged user. Its exact server binary passed local GPU API/chat/vision/parity checks. GPU execution inside Docker could not be exercised on our Docker Desktop host because NVIDIA container support is not configured; this validation limit is separate from the tested native path.
native/: protocol, prefix/branch planner, selected-answer engine and queue bridge.patches/: small, auditable changes to the pinned llama.cpp runtime/server.scripts/: dependency-free build, launch and validation tools.tests/: protocol/planner tests and public synthetic runtime probes.manifests/: artifact identity and release status.Maintainer checks cover automated checks and clean source packaging.
The code is MIT licensed. Model weights retain their own model terms. Nothing is published automatically by these scripts.
Local typed decisions, chat and vision with Winnow-12B and llama.cpp
See the codeA native llama.cpp server for local typed decisions and regular chat, sharing one
loaded model. /v1/systemone evaluates noul, choice, and score questions
against a shared state. /v1/chat/completions retains llama-server's chat, vision,
and streaming interfaces.
Native NVIDIA CUDA and Apple Silicon Metal + Accelerate backends expose the same decision, chat and vision APIs.
This release also contains an explicitly experimental adaptive decision client. Direct decisions remain the default. The optional Linux/CUDA text-only mode uses verified model artifacts and optional matching MTP assistants and frozen model-specific policies; it is not an accuracy-preservation guarantee. See the adaptive contract and limits.
E4B adaptive policy update: the opt-in E4B client now uses the tested
e4b-calibrated75-g95-v1 policy. On 288 text decisions it matched 198 labels
versus 185 with the previous E4B policy, at higher reasoning cost (resident
mean 376 versus 263 ms). The previous E4B policy remains available by ID;
12B Q8 and NVFP4 policies, model weights, and the server binary are unchanged.
Direct decisions remain the default. See results and limits
and upgrade instructions.
64K context and vision on a 16 GB RTX 5070 Ti, with Q8 weights fully on the GPU. Winnow-12B is a fine-tune of Gemma 4 12B; the same loaded model serves typed decisions and regular chat.
Model weights and model card · Benchmarks · API · Installation
On the 231-item public JevBench subset, Winnow-12B Q8 matched hosted Jev at 85.7% (198/231); Winnow-12B BF16 scored 85.3%. Local-model comparisons use the same RTX PRO 5000 Blackwell. See the full benchmark report.

BF16 GGUF, Q8 GGUF, and the matching vision projector are separate downloads. The training dataset is private. Exact release checksums are in the model manifest.
| 12B Q8 method | Correct / 231 | Accuracy |
|---|---|---|
| Native direct | 198 | 85.71% |
| Same-model reasoning, MTP off | 208 | 90.04% |
| Same-model reasoning, MTP draft depth 1 | 209 | 90.48% |
This all-cases, 192-token recipe scored partial text and used no gate or blend;
it is not the released q8-fixed50-v1 policy or its separate 96-decision pilot.
The MTP variant was faster but regressed from 83/96 to 81/96 on an external panel.
Full Q8 benchmark, completion handling, probability losses and timing.
A historical text-only experiment added same-model reasoning below a raw confidence threshold, then blended direct and augmented probabilities equally. It improved Jev and Kev-clean results but reduced Typed teacher agreement (direct → adaptive):
| Historical model / recipe | Jev, 231 | Kev-clean, 1,046 | Typed agreement, 2,000 |
|---|---|---|---|
| 12B NVFP4, fixed gate | 83.55 → 87.88% | 77.82 → 81.45% | 70.60 → 70.15% |
| E4B Q8, fixed gate | 80.52 → 83.55% | 72.66 → 76.96% | 72.35 → 69.50% |
These historical recipes differ from the released NVFP4 entropy and E4B calibrated policies. Full-router means were 347 ms and 234 ms; their included direct HTTP calls averaged 51 ms and 37 ms. Later E4B calibration improved probability losses with little agreement change; the small Q8 confirmation remained inconclusive. See full counts, policy comparisons, latency definitions and MTP limits. Direct decisions remain the default.
Choose direct decisions, reasoning and MTP with one model selector and explicit on/off switches. The guide covers verified downloads, source setup, the thin runtime, and supported combinations for 12B Q8, 12B NVFP4 and E4B Q8.
Supported: Linux + NVIDIA GPU (the measured profile uses a 16 GB RTX 5070 Ti) or Apple Silicon Mac (24 GB or more unified memory for the documented profile). Use a terminal to start the local API server.
git clone https://github.com/EldanRing/winnow-inference.git
cd winnow-inference
Apple Silicon Mac — native Metal + Accelerate:
xcode-select --install
brew install python cmake openssl@3
Wait for Xcode command-line tools to finish installing. The brew command assumes
Homebrew is installed. Setup locates its OpenSSL automatically.
Linux/NVIDIA — native CUDA:
sudo apt-get install build-essential cmake git python3 libssl-dev
Install a compatible NVIDIA driver and the CUDA toolkit. A driver alone is not enough; Blackwell needs CUDA 12.8 or newer. The setup script checks that your toolkit supports your GPU before it downloads anything.
python3 scripts/winnow.py presets
python3 scripts/winnow.py setup --model q8
Short presets are q8, nv4 and e4b. Each defaults to 8K text with MTP and
reasoning off. Setup checks prerequisites, downloads verified weights and builds.
It does not install system packages. Existing valid downloads are reused.
python3 scripts/winnow.py serve --model q8 --context 8k
# In a second terminal:
python3 scripts/winnow.py decide --model q8 --input examples/decisions.json
Select --vision on, --mtp on or --reasoning on explicitly; download the
matching assets first with the same flags. Context and numerical settings can be
overridden where supported. The quickstart includes the mode
matrix, memory estimates, measured baselines and examples. Changed configurations
do not inherit the measured calibration or performance claims.
The server uses http://127.0.0.1:8091. Leave its terminal running; Ctrl+C stops
it. Direct clients verify the chosen model and quantization. Use --model-dir
for another verified asset directory. The existing scripts/serve.py remains
available for custom GGUF paths and legacy platform profiles.
Full installation and troubleshooting ·
API, vision and concurrency · python3 scripts/winnow.py --help
For the default 12B Q8 text-only, 8K setup above, this short release check starts and stops its own authenticated local servers, including selected/full-head comparison. These model paths are Q8-specific; they do not select E4B or NVFP4:
python3 scripts/release_check.py --model models/gguf/Winnow-12B-Q8_0.gguf \
--context 8192 --memory exclusive --output results/release-smoke-text
It uses a 4,096-token output allowance and does not run a quality benchmark or context sweep. For an already running text server, use the individual checks:
python3 scripts/check.py --output results/check
python3 scripts/bench.py --output results/timings --repeats 5
python3 scripts/parity.py --output results/selected
For the optional 12B Q8 64K vision check, first download the projector explicitly. Use a host that meets the 64K platform requirements:
python3 scripts/winnow.py download --model q8 --vision on
python3 scripts/release_check.py --model models/gguf/Winnow-12B-Q8_0.gguf \
--mmproj models/gguf/mmproj-F16.gguf --context 65536 \
--memory exclusive --output results/release-smoke-vision
Add --long-context to that release check explicitly for capacity confirmation.
The following probes require an already running 64K vision server, launched
with the matching platform command in INSTALL:
python3 scripts/check.py --output results/check-vision --vision
python3 scripts/check.py --output results/64k --vision --long-context 65536 --image-size 1536
Authenticated servers are supported through WINNOW_API_KEY_FILE or WINNOW_API_KEY.
Restart with --head full, then use scripts/parity.py --output results/full --compare results/selected to verify selected-head probabilities. The full-head
mode is a reference implementation; normal chat always retains its full head.
All scripts save requests and responses locally. These focused probes validate
runtime behavior and timings, not general model quality or calibrated confidence.
Build the container locally on Linux. GPU execution requires Docker Engine with the NVIDIA Container Toolkit and a CUDA 13.0-compatible driver. Apple Silicon uses native Metal above.
For RTX 50-series / SM120, download both 12B Q8 and its vision projector before starting the default 64K vision container:
python3 scripts/winnow.py download --model q8 --vision on
docker build --build-arg CUDA_ARCH=120 -t winnow-inference .
docker run --rm --gpus 'device=0' -p 127.0.0.1:8091:8091 \
-v "$PWD/models:/models:ro" winnow-inference
The default is the 5070 Ti 64K + vision profile. For a complete container-only
setup, including downloading weights without host Python, follow
the container quickstart.
Other NVIDIA architectures can build the Dockerfile with their own CUDA_ARCH.
The runtime image is about 1.7 GB uncompressed, excludes model weights and build tools, and runs as an unprivileged user. Its exact server binary passed local GPU API/chat/vision/parity checks. GPU execution inside Docker could not be exercised on our Docker Desktop host because NVIDIA container support is not configured; this validation limit is separate from the tested native path.
native/: protocol, prefix/branch planner, selected-answer engine and queue bridge.patches/: small, auditable changes to the pinned llama.cpp runtime/server.scripts/: dependency-free build, launch and validation tools.tests/: protocol/planner tests and public synthetic runtime probes.manifests/: artifact identity and release status.Maintainer checks cover automated checks and clean source packaging.
The code is MIT licensed. Model weights retain their own model terms. Nothing is published automatically by these scripts.