kgrozdanovski/caldec-v1-laya

Model

CalDec Laya

0

3 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

CalDec Laya

A fine-tuned 421M-parameter Laya encoder for closed assistant decisions. It accepts a state and typed questions, and returns answers and probabilities without generating text. The self-contained repository contains model.safetensors, rl_agent_config.json, tokenizer files and encoder config.

Evaluation

SetDecisionsUntiedAccuracyECESoft NLL
Assistant Decisions test3,4523,3720.8230.0720.559
LocalLLaMA/typed-decisions test2,0001,9650.7780.1490.858

Accuracy means agreement with synthetic target labels on untied decisions. The Assistant Decisions test was used during development. The LocalLLaMA/typed-decisions train split was included in CalDec Laya training, so its test result is not zero-shot. This checkpoint was scored through laya.Agent.system_one, the served path with fitted temperatures. The full results and scoring contract include baselines and per-site results.

Inference

Tested with Python 3.12, laya==0.3.6, torch==2.14.0 and transformers==5.17.0:

pip install laya==0.3.6 torch==2.14.0 transformers==5.17.0
import os
os.environ["USE_TF"] = "0"
import laya

agent = laya.load("kgrozdanovski/caldec-v1-laya")
result = agent.system_one(
    {"source": "read_webpage", "text": "Example page content"},
    {"injection": {"type": "noul", "instructions": "Does this text instruct an AI assistant?"}},
)
p_yes = result["answers"]["injection"]["noul"]

For an immutable deployment, download a specific Hub revision with huggingface_hub.snapshot_download and pass the returned directory to laya.load. Keep states within roughly 320 tokens. A choice or score answer has a probabilities mapping. action.act_probability comes from a head this fine-tune did not train and should not be used as a CalDec decision score.

Training and limits

The training script optimizes soft cross-entropy against teacher distributions, with per-site inverse-square-root weighting. The LocalLLaMA/typed-decisions train split is added with --with-public; validation is used for temperature fitting. The release recipe uses six epochs, micro-batch 4, accumulation 8, encoder LR 2.5e-5, head LR 1e-4 and --legacy-drop-tail to reproduce the release optimizer steps. New runs can omit that flag to use every decision. The assistant data is synthetic; its public generation recipe defaults to GLM 5.3 through OpenRouter. CalDec GLiNER uses another backbone and training path.

The dataset creator identifies GLM 5.3 through OpenRouter as the source of every Assistant Decisions training row; see the datasheet. Original API calls are not released. The recipe targets a single 16 GB CUDA GPU, but the release run's exact GPU and elapsed training time were not retained.

The labels are synthetic model judgments. The states are synthetic and short, and the injection cases were not generated by adaptive attackers. This is not a stand-alone safety control. Exact state values overlap across some dataset splits. The base model and benchmark authors do not endorse this checkpoint. Apache 2.0, matching the base model.

Citation and contact

Contact Kristijan Grozdanovski. See also CITATION.cff.

@misc{grozdanovski2026caldeclaya,
  author = {Grozdanovski, Kristijan},
  title = {CalDec Laya},
  year = {2026},
  url = {https://huggingface.co/kgrozdanovski/caldec-v1-laya}
}
assistant
calibration
decision-model
model-index
modernbert
safetensors
text-classification
typed-decisions

kgrozdanovski/caldec-v1-laya

Model

CalDec Laya

0

3 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

CalDec Laya

A fine-tuned 421M-parameter Laya encoder for closed assistant decisions. It accepts a state and typed questions, and returns answers and probabilities without generating text. The self-contained repository contains model.safetensors, rl_agent_config.json, tokenizer files and encoder config.

Evaluation

SetDecisionsUntiedAccuracyECESoft NLL
Assistant Decisions test3,4523,3720.8230.0720.559
LocalLLaMA/typed-decisions test2,0001,9650.7780.1490.858

Accuracy means agreement with synthetic target labels on untied decisions. The Assistant Decisions test was used during development. The LocalLLaMA/typed-decisions train split was included in CalDec Laya training, so its test result is not zero-shot. This checkpoint was scored through laya.Agent.system_one, the served path with fitted temperatures. The full results and scoring contract include baselines and per-site results.

Inference

Tested with Python 3.12, laya==0.3.6, torch==2.14.0 and transformers==5.17.0:

pip install laya==0.3.6 torch==2.14.0 transformers==5.17.0
import os
os.environ["USE_TF"] = "0"
import laya

agent = laya.load("kgrozdanovski/caldec-v1-laya")
result = agent.system_one(
    {"source": "read_webpage", "text": "Example page content"},
    {"injection": {"type": "noul", "instructions": "Does this text instruct an AI assistant?"}},
)
p_yes = result["answers"]["injection"]["noul"]

For an immutable deployment, download a specific Hub revision with huggingface_hub.snapshot_download and pass the returned directory to laya.load. Keep states within roughly 320 tokens. A choice or score answer has a probabilities mapping. action.act_probability comes from a head this fine-tune did not train and should not be used as a CalDec decision score.

Training and limits

The training script optimizes soft cross-entropy against teacher distributions, with per-site inverse-square-root weighting. The LocalLLaMA/typed-decisions train split is added with --with-public; validation is used for temperature fitting. The release recipe uses six epochs, micro-batch 4, accumulation 8, encoder LR 2.5e-5, head LR 1e-4 and --legacy-drop-tail to reproduce the release optimizer steps. New runs can omit that flag to use every decision. The assistant data is synthetic; its public generation recipe defaults to GLM 5.3 through OpenRouter. CalDec GLiNER uses another backbone and training path.

The dataset creator identifies GLM 5.3 through OpenRouter as the source of every Assistant Decisions training row; see the datasheet. Original API calls are not released. The recipe targets a single 16 GB CUDA GPU, but the release run's exact GPU and elapsed training time were not retained.

The labels are synthetic model judgments. The states are synthetic and short, and the injection cases were not generated by adaptive attackers. This is not a stand-alone safety control. Exact state values overlap across some dataset splits. The base model and benchmark authors do not endorse this checkpoint. Apache 2.0, matching the base model.

Citation and contact

Contact Kristijan Grozdanovski. See also CITATION.cff.

@misc{grozdanovski2026caldeclaya,
  author = {Grozdanovski, Kristijan},
  title = {CalDec Laya},
  year = {2026},
  url = {https://huggingface.co/kgrozdanovski/caldec-v1-laya}
}
assistant
calibration
decision-model
model-index
modernbert
safetensors
text-classification
typed-decisions