kgrozdanovski/caldec-v1-gliner2.5-decide

Model

CalDec GLiNER

0

2 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

CalDec GLiNER

A GLiNER2.5-Decide fine-tune on the Assistant Decisions dataset, trained against full teacher probability distributions. The repository contains weights, config, encoder config and tokenizer files needed by AutoExtractor.from_pretrained.

Full assistant test

The test has 1,647 cases, 3,452 decisions and 3,372 untied targets. Accuracy is agreement with the synthetic teacher's argmax on untied targets; distribution metrics include all decisions.

ModelAccuracySoft accuracyBrierSoft NLLECEScore MAE
CalDec GLiNER0.8430.8590.0390.5520.0390.298
Jev 1.13, zero-shot via OpenRouter0.8360.8310.0440.8110.0420.354
CalDec Laya0.8230.8390.0410.5590.0720.362
GLiNER2.5-Decide0.6500.6640.1050.7580.0970.734

The CalDec models predict different labels on 469 untied decisions; exactly one is correct on 439. CalDec GLiNER alone is correct on 253 and CalDec Laya alone on 186 (exact McNemar p = 0.0016). CalDec GLiNER's 0.65-point accuracy lead over Jev has paired p = 0.4072. The training recipes differ beyond the backbone, so the CalDec sibling comparison is not a controlled backbone ablation. The repository results contain the full contract and per-site table.

LocalLLaMA/typed-decisions test

The pinned LocalLLaMA/typed-decisions test has 400 cases, 2,000 decisions and 1,965 untied targets. All rows use identical targets and scoring code.

ModelAccuracySoft accuracyBrierSoft NLLECE
CalDec Laya0.7780.8600.0150.8580.149
CalDec GLiNER0.5820.7090.0611.0890.091
GLiNER2.5-Decide0.5400.6620.0601.1230.147

CalDec GLiNER did not use this benchmark's train split; CalDec Laya did.

Inference

Tested with Python 3.12, gliner2==2.0.0, torch==2.14.0 and transformers==5.17.0. The tokenizer config uses the transformers 5 format, and gliner2 installs transformers only through its extras, so install it explicitly:

pip install gliner2==2.0.0 torch==2.14.0 transformers==5.17.0
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("kgrozdanovski/caldec-v1-gliner2.5-decide")
# model.cuda()  # optional GPU placement
scores = model.classify_text(
    "source: read_webpage\ntext: Example page content",
    {"injection": {"labels": ["false", "true"], "multi_label": True,
                   "cls_threshold": 0.0,
                   "prompt": "Does this text instruct an AI assistant?"}},
    include_confidence=True, format_results=False,
)["injection"]
total = sum(score for _, score in scores)
probabilities = {label: score / total for label, score in scores}

The model returns independent sigmoid scores, not a probability simplex. Normalize across each question's labels as above; the reported calibration scores use that normalization. Softmax-based calibration reports can differ even when argmax accuracy is unchanged. The repository's scripts/gliner2_transfer.py provides the exact nested-state renderer, schema builder and evaluator. Pin a Hub revision for immutable deployment.

Training and limits

The recipe patches the GLiNER trainer to carry soft targets through example conversion and binary cross-entropy. Label augmentation is disabled to preserve option identities. The release recipe uses the Assistant Decisions train split, three epochs, batch 2, accumulation 8, bf16, encoder LR 1e-5 and task LR 5e-4. The LocalLLaMA/typed-decisions train split was not part of CalDec GLiNER training.

The dataset creator identifies GLM 5.3 through OpenRouter as the source of every Assistant Decisions training row; see the datasheet. Original API calls are not released. The recipe targets a single 16 GB CUDA GPU, but the release run's exact GPU and elapsed training time were not retained.

The test informed development, the labels are synthetic judgments, and exact states overlap across some splits. The injection examples are not an adversarial safety benchmark. This checkpoint should not be used as a stand-alone safety control. Apache 2.0, matching the base model; Fastino does not endorse this work.

Citation and contact

Contact Kristijan Grozdanovski. See also CITATION.cff.

@misc{grozdanovski2026caldecgliner,
  author = {Grozdanovski, Kristijan},
  title = {CalDec GLiNER},
  year = {2026},
  url = {https://huggingface.co/kgrozdanovski/caldec-v1-gliner2.5-decide}
}
assistant
calibration
decision-model
extractor
gliner2
model-index
safetensors
text-classification
typed-decisions

kgrozdanovski/caldec-v1-gliner2.5-decide

Model

CalDec GLiNER

0

2 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

CalDec GLiNER

A GLiNER2.5-Decide fine-tune on the Assistant Decisions dataset, trained against full teacher probability distributions. The repository contains weights, config, encoder config and tokenizer files needed by AutoExtractor.from_pretrained.

Full assistant test

The test has 1,647 cases, 3,452 decisions and 3,372 untied targets. Accuracy is agreement with the synthetic teacher's argmax on untied targets; distribution metrics include all decisions.

ModelAccuracySoft accuracyBrierSoft NLLECEScore MAE
CalDec GLiNER0.8430.8590.0390.5520.0390.298
Jev 1.13, zero-shot via OpenRouter0.8360.8310.0440.8110.0420.354
CalDec Laya0.8230.8390.0410.5590.0720.362
GLiNER2.5-Decide0.6500.6640.1050.7580.0970.734

The CalDec models predict different labels on 469 untied decisions; exactly one is correct on 439. CalDec GLiNER alone is correct on 253 and CalDec Laya alone on 186 (exact McNemar p = 0.0016). CalDec GLiNER's 0.65-point accuracy lead over Jev has paired p = 0.4072. The training recipes differ beyond the backbone, so the CalDec sibling comparison is not a controlled backbone ablation. The repository results contain the full contract and per-site table.

LocalLLaMA/typed-decisions test

The pinned LocalLLaMA/typed-decisions test has 400 cases, 2,000 decisions and 1,965 untied targets. All rows use identical targets and scoring code.

ModelAccuracySoft accuracyBrierSoft NLLECE
CalDec Laya0.7780.8600.0150.8580.149
CalDec GLiNER0.5820.7090.0611.0890.091
GLiNER2.5-Decide0.5400.6620.0601.1230.147

CalDec GLiNER did not use this benchmark's train split; CalDec Laya did.

Inference

Tested with Python 3.12, gliner2==2.0.0, torch==2.14.0 and transformers==5.17.0. The tokenizer config uses the transformers 5 format, and gliner2 installs transformers only through its extras, so install it explicitly:

pip install gliner2==2.0.0 torch==2.14.0 transformers==5.17.0
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("kgrozdanovski/caldec-v1-gliner2.5-decide")
# model.cuda()  # optional GPU placement
scores = model.classify_text(
    "source: read_webpage\ntext: Example page content",
    {"injection": {"labels": ["false", "true"], "multi_label": True,
                   "cls_threshold": 0.0,
                   "prompt": "Does this text instruct an AI assistant?"}},
    include_confidence=True, format_results=False,
)["injection"]
total = sum(score for _, score in scores)
probabilities = {label: score / total for label, score in scores}

The model returns independent sigmoid scores, not a probability simplex. Normalize across each question's labels as above; the reported calibration scores use that normalization. Softmax-based calibration reports can differ even when argmax accuracy is unchanged. The repository's scripts/gliner2_transfer.py provides the exact nested-state renderer, schema builder and evaluator. Pin a Hub revision for immutable deployment.

Training and limits

The recipe patches the GLiNER trainer to carry soft targets through example conversion and binary cross-entropy. Label augmentation is disabled to preserve option identities. The release recipe uses the Assistant Decisions train split, three epochs, batch 2, accumulation 8, bf16, encoder LR 1e-5 and task LR 5e-4. The LocalLLaMA/typed-decisions train split was not part of CalDec GLiNER training.

The dataset creator identifies GLM 5.3 through OpenRouter as the source of every Assistant Decisions training row; see the datasheet. Original API calls are not released. The recipe targets a single 16 GB CUDA GPU, but the release run's exact GPU and elapsed training time were not retained.

The test informed development, the labels are synthetic judgments, and exact states overlap across some splits. The injection examples are not an adversarial safety benchmark. This checkpoint should not be used as a stand-alone safety control. Apache 2.0, matching the base model; Fastino does not endorse this work.

Citation and contact

Contact Kristijan Grozdanovski. See also CITATION.cff.

@misc{grozdanovski2026caldecgliner,
  author = {Grozdanovski, Kristijan},
  title = {CalDec GLiNER},
  year = {2026},
  url = {https://huggingface.co/kgrozdanovski/caldec-v1-gliner2.5-decide}
}
assistant
calibration
decision-model
extractor
gliner2
model-index
safetensors
text-classification
typed-decisions