A Q4_K_M quantization of EldanRing/Winnow-12B, a Gemma 4 12B IT fine-tune for typed decisions. The upstream repository ships BF16 and Q8_0. This file fits a 12 GB GPU with full offload: Ollama loads it in 8.0 GB.
| File | Size | SHA-256 |
|---|---|---|
Winnow-12B-Q4_K_M.gguf | 7.38 GB | 90c1bcc54d26bc99207929d4be23b9a88dc3e944fe39e80ce8bf95225d516218 |
Text only. Vision with the upstream mmproj-Winnow-12B.gguf was not tested with this quantization.
gguf/Winnow-12B-BF16.gguf from EldanRing/Winnow-12B, SHA-256 10e6b41e00a3fc6668b53a92fb5d9c85179f47145e0786e7a3bc7b5b09ba8968.llama-quantize Winnow-12B-BF16.gguf Winnow-12B-Q4_K_M.gguf Q4_K_M, built unpatched from llama.cpp commit 911f6cd, the commit Winnow's inference server pins.printf 'FROM hf.co/Piotr1215/Winnow-12B-GGUF:Q4_K_M\nTEMPLATE {{ .Prompt }}\nRENDERER gemma4\nPARSER gemma4\nPARAMETER temperature 1\nPARAMETER top_k 64\nPARAMETER top_p 0.95\n' > Modelfile
ollama create winnow:12b-q4_K_M -f Modelfile
The Modelfile selects Ollama's built-in Gemma 4 renderer, the setup measured below. To read a decision, ask for one token with log probabilities (num_predict: 1, logprobs: true, top_logprobs: 20, think: false) and compare the answer labels' probabilities.
Winnow's own llama.cpp server adds a typed-decision API. It was not tested with this quantization.
One comparison against stock gemma4:12b, both Q4_K_M, served by Ollama 0.34.3 with the same Gemma 4 renderer and the same prompts, on an RTX 5070 Ti Laptop GPU (12 GB). Each answer is read from the first token's log probabilities.
| Set | Items | gemma4:12b | This quantization |
|---|---|---|---|
| SemIf authored decisions, three described options | 144 | 139 | 139 |
| Synthetic yes/no | 15 | 15 | 15 |
| Private RAG relevance | 70 | 66 | 66 |
| Private mail, newsletter or not | 55 | 55 | 55 |
| Private mail, needs action or not | 60 | 50 | 54 |
| Total | 344 | 325 | 329 |
| Measure | gemma4:12b | This quantization |
|---|---|---|
| Held-out NLL, yes/no sets, fitted temperature | 0.200 | 0.166 |
| Held-out NLL, option sets, fitted temperature | 0.171 | 0.167 |
| Model-call latency p50 / p95 | 218 / 601 ms | 222 / 622 ms |
The total differs by 4 items: 7 won, 3 lost. A paired bootstrap puts the 95% interval for the accuracy difference at -0.6 to +2.9 points, so treat it as parity with a gain on the action set. This quantization was not compared with the upstream BF16 or Q8_0 files, so its quantization loss is unmeasured.
Apache-2.0, following Winnow-12B and Gemma 4. LICENSE and NOTICE are copied unchanged from EldanRing/Winnow-12B. The only change from the upstream BF16 file is the Q4_K_M quantization described above. This repository is independent and not affiliated with or endorsed by EldanRing, TypeSafe or Google.
A Q4_K_M quantization of EldanRing/Winnow-12B, a Gemma 4 12B IT fine-tune for typed decisions. The upstream repository ships BF16 and Q8_0. This file fits a 12 GB GPU with full offload: Ollama loads it in 8.0 GB.
| File | Size | SHA-256 |
|---|---|---|
Winnow-12B-Q4_K_M.gguf | 7.38 GB | 90c1bcc54d26bc99207929d4be23b9a88dc3e944fe39e80ce8bf95225d516218 |
Text only. Vision with the upstream mmproj-Winnow-12B.gguf was not tested with this quantization.
gguf/Winnow-12B-BF16.gguf from EldanRing/Winnow-12B, SHA-256 10e6b41e00a3fc6668b53a92fb5d9c85179f47145e0786e7a3bc7b5b09ba8968.llama-quantize Winnow-12B-BF16.gguf Winnow-12B-Q4_K_M.gguf Q4_K_M, built unpatched from llama.cpp commit 911f6cd, the commit Winnow's inference server pins.printf 'FROM hf.co/Piotr1215/Winnow-12B-GGUF:Q4_K_M\nTEMPLATE {{ .Prompt }}\nRENDERER gemma4\nPARSER gemma4\nPARAMETER temperature 1\nPARAMETER top_k 64\nPARAMETER top_p 0.95\n' > Modelfile
ollama create winnow:12b-q4_K_M -f Modelfile
The Modelfile selects Ollama's built-in Gemma 4 renderer, the setup measured below. To read a decision, ask for one token with log probabilities (num_predict: 1, logprobs: true, top_logprobs: 20, think: false) and compare the answer labels' probabilities.
Winnow's own llama.cpp server adds a typed-decision API. It was not tested with this quantization.
One comparison against stock gemma4:12b, both Q4_K_M, served by Ollama 0.34.3 with the same Gemma 4 renderer and the same prompts, on an RTX 5070 Ti Laptop GPU (12 GB). Each answer is read from the first token's log probabilities.
| Set | Items | gemma4:12b | This quantization |
|---|---|---|---|
| SemIf authored decisions, three described options | 144 | 139 | 139 |
| Synthetic yes/no | 15 | 15 | 15 |
| Private RAG relevance | 70 | 66 | 66 |
| Private mail, newsletter or not | 55 | 55 | 55 |
| Private mail, needs action or not | 60 | 50 | 54 |
| Total | 344 | 325 | 329 |
| Measure | gemma4:12b | This quantization |
|---|---|---|
| Held-out NLL, yes/no sets, fitted temperature | 0.200 | 0.166 |
| Held-out NLL, option sets, fitted temperature | 0.171 | 0.167 |
| Model-call latency p50 / p95 | 218 / 601 ms | 222 / 622 ms |
The total differs by 4 items: 7 won, 3 lost. A paired bootstrap puts the 95% interval for the accuracy difference at -0.6 to +2.9 points, so treat it as parity with a gain on the action set. This quantization was not compared with the upstream BF16 or Q8_0 files, so its quantization loss is unmeasured.
Apache-2.0, following Winnow-12B and Gemma 4. LICENSE and NOTICE are copied unchanged from EldanRing/Winnow-12B. The only change from the upstream BF16 file is the Q4_K_M quantization described above. This repository is independent and not affiliated with or endorsed by EldanRing, TypeSafe or Google.