Local decision model with calibrated probabilities: send a state and yes/no, choice or score questions, get a probability for every option. 0.8B GGUF on CPU, Jev-style API.
Python
2
1 commits
updated Oct 1, 2026
Open, local decision models with calibrated probabilities. Send a state (text or JSON) and typed questions (yes/no, choice, score); get a probability for every option back. No text generation, no network calls, no per-call cost. The API follows the Jev-style Decisions format, so existing clients mostly need only a new base URL.
gutsy-0.8b v0.3 is a fine-tune of Qwen3.5-0.8B, served as a 775 MB GGUF by llama.cpp on an
ordinary CPU.
| Model | gutsy-0.8b v0.3 (Q8_0 GGUF, 775 MB): kouhxp/gutsy |
| Runtime | gutsy-inference (this repository) |
| Model card | MODEL_CARD.md |
| License | Apache 2.0 (see Licensing) |
reject probability for "none of
the options fits, or the state doesn't say".git clone https://github.com/kouhxp/gutsy
pip install -e gutsy/gutsy-inference # needs llama-cpp-python >= 0.3.35
pip install -U huggingface_hub
hf download kouhxp/gutsy gutsy-0.8b-v03-q8_0.gguf gutsy-0.8b-v03.calibration.json \
--local-dir gutsy/gutsy-inference/models
cd gutsy/gutsy-inference
cp models.example.json models.json
gutsy-inference check # expect "caching OK"
gutsy-inference serve # http://127.0.0.1:8765
curl -s http://127.0.0.1:8765/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "gutsy-0.8b-v03",
"state": "Order 1182 shipped Monday. The customer says it arrived Thursday with the box crushed.",
"questions": {"decision": {"type": "choice", "instructions": "What should support do next?",
"criteria": {"refund": "issue a refund", "replace": "send a replacement",
"escalate": "a manager reviews it within 24 hours"}}}}'
Endpoints: POST /v1/systemone (TypeSafe-style; works with JevBench's typesafe adapter),
POST /api/alpha/decisions and /v1/decisions (OpenRouter Decisions style), GET /health.
Each answer has probabilities; choice and score answers add confidence (how concentrated the
distribution is, from 0 for uniform to 1), margin (top two apart) and reject. Gate risky
actions on thresholds you measure on your own labeled data; see the
runtime README.
All numbers below were run by us. v0.1 and v0.2 are earlier internal versions (never released), shown for context. Comparisons with other systems are indicative: different harness versions and item samples, and the other systems' numbers are self-reported unless noted.
| tier | v0.1 | v0.2 | v0.3 |
|---|---|---|---|
| easy (48) | 45 | 48 | 47 |
| original (72) | 46 | 46 | 61 |
| hard (111) | 49 | 41 | 59 |
| total | 140 (0.606) | 135 (0.584) | 167 (0.723) |
| calibration error (hard) | 0.167 | 0.254 | 0.078 |
| ordinal MAE (original / hard) | 0.53 / 0.73 | 0.55 / 0.46 | 0.30 / 0.59 |
Reported by other projects on the same public set: Jev 0.866, decision-4b 0.883, Neriv 0.6B 0.636, XERON-0.4 0.558, tuned Laya 0.537; Jeff-0.8B 47.6% on the hard tier (105 items). JevBench describes public-set results as preliminary; official results use sealed items.
Hard-tier families for v0.3 (5-19 items each, so single families are noisy): routing_hard 1.00, trap 0.88, adversarial 0.83, tradeoff 0.67, judge_hard 0.65, multi_hop 0.50, probability 0.50, long_policy 0.42, temporal_numeric 0.33, ambiguous 0.00.
| Qwen3.5-0.8B (no fine-tune) | gutsy v0.1 | gutsy v0.3 | Jev (typesafe/jev-1.13, via API) | |
|---|---|---|---|---|
| BoolQ | 0.700 | 0.821 | 0.850 | 0.906 |
| SNLI | 0.489 | 0.886 | 0.879 | 0.883 |
| CommonsenseQA (never trained on) | 0.463 | 0.476 | 0.587 | 0.863 |
Jev was queried through its API for benchmarking only; none of its output was used for training.
| v0.1 | v0.2 | v0.3 | Jeff-0.8B v1.0 | Jev | |
|---|---|---|---|---|---|
| BBH | 0.419 | 0.413 | 0.464 | 64.0 | 94.3 |
| Financial PhraseBank | 0.857 | 0.481 | 0.921 | 96.4 | 77.0 |
| JudgeBench | 0.416 | 0.424 | 0.596 | 62.6 | 78.6 |
| RAGTruth † | 0.469 | 0.646 | 0.790 | 86.1 | 77.3 |
| WinoGrande † | 0.503 | 0.516 | 0.657 | 68.6 | 90.7 |
† In-domain for v0.3: the RAGTruth and WinoGrande training splits are in its training data (test items never are). Financial PhraseBank is close to the financial-tweet sentiment data used in training. Jeff and Jev figures are from Jeff's README, on a different sample of these benchmarks.
Don't automate on these without your own testing:
Full details, sources and licenses: MODEL_CARD.md.
Code and model weights: Apache 2.0 (see LICENSE). The base model (Qwen3.5-0.8B) and the
teacher model (Qwen3.8-27B) are Apache 2.0. The training data draws on CC BY 4.0, CC BY-SA, CC0,
MIT and Apache 2.0 sources, credited in MODEL_CARD.md; Jev Decisions v1 is CC BY 4.0 with
upstream NVIDIA terms. Users who redistribute the training data itself must follow each source's
terms.
Python
100.0%
Local decision model with calibrated probabilities: send a state and yes/no, choice or score questions, get a probability for every option. 0.8B GGUF on CPU, Jev-style API.
Python
2
1 commits
updated Oct 1, 2026
Open, local decision models with calibrated probabilities. Send a state (text or JSON) and typed questions (yes/no, choice, score); get a probability for every option back. No text generation, no network calls, no per-call cost. The API follows the Jev-style Decisions format, so existing clients mostly need only a new base URL.
gutsy-0.8b v0.3 is a fine-tune of Qwen3.5-0.8B, served as a 775 MB GGUF by llama.cpp on an
ordinary CPU.
| Model | gutsy-0.8b v0.3 (Q8_0 GGUF, 775 MB): kouhxp/gutsy |
| Runtime | gutsy-inference (this repository) |
| Model card | MODEL_CARD.md |
| License | Apache 2.0 (see Licensing) |
reject probability for "none of
the options fits, or the state doesn't say".git clone https://github.com/kouhxp/gutsy
pip install -e gutsy/gutsy-inference # needs llama-cpp-python >= 0.3.35
pip install -U huggingface_hub
hf download kouhxp/gutsy gutsy-0.8b-v03-q8_0.gguf gutsy-0.8b-v03.calibration.json \
--local-dir gutsy/gutsy-inference/models
cd gutsy/gutsy-inference
cp models.example.json models.json
gutsy-inference check # expect "caching OK"
gutsy-inference serve # http://127.0.0.1:8765
curl -s http://127.0.0.1:8765/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "gutsy-0.8b-v03",
"state": "Order 1182 shipped Monday. The customer says it arrived Thursday with the box crushed.",
"questions": {"decision": {"type": "choice", "instructions": "What should support do next?",
"criteria": {"refund": "issue a refund", "replace": "send a replacement",
"escalate": "a manager reviews it within 24 hours"}}}}'
Endpoints: POST /v1/systemone (TypeSafe-style; works with JevBench's typesafe adapter),
POST /api/alpha/decisions and /v1/decisions (OpenRouter Decisions style), GET /health.
Each answer has probabilities; choice and score answers add confidence (how concentrated the
distribution is, from 0 for uniform to 1), margin (top two apart) and reject. Gate risky
actions on thresholds you measure on your own labeled data; see the
runtime README.
All numbers below were run by us. v0.1 and v0.2 are earlier internal versions (never released), shown for context. Comparisons with other systems are indicative: different harness versions and item samples, and the other systems' numbers are self-reported unless noted.
| tier | v0.1 | v0.2 | v0.3 |
|---|---|---|---|
| easy (48) | 45 | 48 | 47 |
| original (72) | 46 | 46 | 61 |
| hard (111) | 49 | 41 | 59 |
| total | 140 (0.606) | 135 (0.584) | 167 (0.723) |
| calibration error (hard) | 0.167 | 0.254 | 0.078 |
| ordinal MAE (original / hard) | 0.53 / 0.73 | 0.55 / 0.46 | 0.30 / 0.59 |
Reported by other projects on the same public set: Jev 0.866, decision-4b 0.883, Neriv 0.6B 0.636, XERON-0.4 0.558, tuned Laya 0.537; Jeff-0.8B 47.6% on the hard tier (105 items). JevBench describes public-set results as preliminary; official results use sealed items.
Hard-tier families for v0.3 (5-19 items each, so single families are noisy): routing_hard 1.00, trap 0.88, adversarial 0.83, tradeoff 0.67, judge_hard 0.65, multi_hop 0.50, probability 0.50, long_policy 0.42, temporal_numeric 0.33, ambiguous 0.00.
| Qwen3.5-0.8B (no fine-tune) | gutsy v0.1 | gutsy v0.3 | Jev (typesafe/jev-1.13, via API) | |
|---|---|---|---|---|
| BoolQ | 0.700 | 0.821 | 0.850 | 0.906 |
| SNLI | 0.489 | 0.886 | 0.879 | 0.883 |
| CommonsenseQA (never trained on) | 0.463 | 0.476 | 0.587 | 0.863 |
Jev was queried through its API for benchmarking only; none of its output was used for training.
| v0.1 | v0.2 | v0.3 | Jeff-0.8B v1.0 | Jev | |
|---|---|---|---|---|---|
| BBH | 0.419 | 0.413 | 0.464 | 64.0 | 94.3 |
| Financial PhraseBank | 0.857 | 0.481 | 0.921 | 96.4 | 77.0 |
| JudgeBench | 0.416 | 0.424 | 0.596 | 62.6 | 78.6 |
| RAGTruth † | 0.469 | 0.646 | 0.790 | 86.1 | 77.3 |
| WinoGrande † | 0.503 | 0.516 | 0.657 | 68.6 | 90.7 |
† In-domain for v0.3: the RAGTruth and WinoGrande training splits are in its training data (test items never are). Financial PhraseBank is close to the financial-tweet sentiment data used in training. Jeff and Jev figures are from Jeff's README, on a different sample of these benchmarks.
Don't automate on these without your own testing:
Full details, sources and licenses: MODEL_CARD.md.
Code and model weights: Apache 2.0 (see LICENSE). The base model (Qwen3.5-0.8B) and the
teacher model (Qwen3.8-27B) are Apache 2.0. The training data draws on CC BY 4.0, CC BY-SA, CC0,
MIT and Apache 2.0 sources, credited in MODEL_CARD.md; Jev Decisions v1 is CC BY 4.0 with
upstream NVIDIA terms. Users who redistribute the training data itself must follow each source's
terms.
Python
100.0%