kgrozdanovski/assistant-decisions

Dataset

Assistant Decisions

0

3 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

Assistant Decisions

7,436 synthetic cases and 15,471 labelled decisions across 18 assistant decision sites. Targets are probability distributions for noul (yes/no), choice and score questions. This dataset supports the CalDec training and evaluation recipes.

SplitCasesDecisions
Train4,98410,375
Validation8051,644
Test1,6473,452

Each row has six string fields: case_id, site, workflow, state, questions, and gold. Decode the last three from JSON text:

import json
from datasets import load_dataset
row = load_dataset("kgrozdanovski/assistant-decisions")["train"][0]
state = json.loads(row["state"])
questions = json.loads(row["questions"])
gold = json.loads(row["gold"])

sites.json gives the 18 site definitions; prompts.json contains site generation and labelling prompts. The dataset creator identifies GLM 5.3 through OpenRouter as the source of every released row, including generated states and target distributions. The exp-, gen- and fill1- case-ID prefixes identify generation builds, not providers. Original API calls and response journals are not released, so individual calls cannot be verified from the public files. The public recipe defaults to GLM 5.3 and shows how to configure other models and OpenAI-compatible providers for new builds. The datasheet describes collection and limitations.

Model comparisons

Assistant accuracy is agreement with the synthetic target's argmax on 3,372 untied test decisions; it is not independent correctness. ECE also excludes ties; distribution metrics include all 3,452 assistant decisions.

ModelAssistant accuracyECESoft NLLLocalLLaMA/typed-decisions test accuracy
CalDec GLiNER0.8430.0390.5520.582
Jev 1.13, zero-shot via OpenRouter0.8360.0420.8110.738
CalDec Laya0.8230.0720.5590.778†
GLiNER2.5-Decide0.6500.0970.7580.540
Laya specialist0.5780.0410.8050.773†
Majority class, fitted on train0.574———
Laya base0.5570.1580.9890.361

† Trained on the LocalLLaMA/typed-decisions train split, so this test score is not zero-shot.

Laya base means the upstream convaiinnovations/laya checkpoint. Laya specialist means the upstream convaiinnovations/laya-typed-decisions checkpoint, fine-tuned on the Typed Decisions train split.

The final column uses the pinned test split of Typed Decisions (LocalLLaMA/typed-decisions), with 2,000 decisions and 1,965 untied targets. CalDec Laya used that benchmark's train split; CalDec GLiNER did not. The full comparison is in RESULTS.md.

The test split was used during development. Three exact state values occur in more than one split; the repository's scripts/check_data.py prints the case ids. This limits test independence slightly. The full evaluation contract and per-site table are in RESULTS.md.

Use and limits

The states are synthetic, English, and generally short. Some contain prompt-injection text and invented credentials by design. Treat all state text as untrusted. These are generated cases, not an adversarial safety benchmark. Targets reflect model judgments, including their errors. Evaluate a trained model on real, independently labelled cases before consequential use.

The repository's datasheet documents collection, format, processing, split overlap and intended use. The license is Apache 2.0. Review the hosted model providers' terms when using the recipe to create new data.

Citation and contact

Contact Kristijan Grozdanovski. The repository's CITATION.cff is the machine-readable citation.

@misc{grozdanovski2026assistantdecisions,
  author = {Grozdanovski, Kristijan},
  title = {Assistant Decisions: A Dataset and Recipe for Small Typed-Decision Models},
  year = {2026},
  url = {https://huggingface.co/datasets/kgrozdanovski/assistant-decisions}
}
assistant
calibration
decision-model
synthetic
typed-decisions

kgrozdanovski/assistant-decisions

Dataset

Assistant Decisions

0

3 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

Assistant Decisions

7,436 synthetic cases and 15,471 labelled decisions across 18 assistant decision sites. Targets are probability distributions for noul (yes/no), choice and score questions. This dataset supports the CalDec training and evaluation recipes.

SplitCasesDecisions
Train4,98410,375
Validation8051,644
Test1,6473,452

Each row has six string fields: case_id, site, workflow, state, questions, and gold. Decode the last three from JSON text:

import json
from datasets import load_dataset
row = load_dataset("kgrozdanovski/assistant-decisions")["train"][0]
state = json.loads(row["state"])
questions = json.loads(row["questions"])
gold = json.loads(row["gold"])

sites.json gives the 18 site definitions; prompts.json contains site generation and labelling prompts. The dataset creator identifies GLM 5.3 through OpenRouter as the source of every released row, including generated states and target distributions. The exp-, gen- and fill1- case-ID prefixes identify generation builds, not providers. Original API calls and response journals are not released, so individual calls cannot be verified from the public files. The public recipe defaults to GLM 5.3 and shows how to configure other models and OpenAI-compatible providers for new builds. The datasheet describes collection and limitations.

Model comparisons

Assistant accuracy is agreement with the synthetic target's argmax on 3,372 untied test decisions; it is not independent correctness. ECE also excludes ties; distribution metrics include all 3,452 assistant decisions.

ModelAssistant accuracyECESoft NLLLocalLLaMA/typed-decisions test accuracy
CalDec GLiNER0.8430.0390.5520.582
Jev 1.13, zero-shot via OpenRouter0.8360.0420.8110.738
CalDec Laya0.8230.0720.5590.778†
GLiNER2.5-Decide0.6500.0970.7580.540
Laya specialist0.5780.0410.8050.773†
Majority class, fitted on train0.574———
Laya base0.5570.1580.9890.361

† Trained on the LocalLLaMA/typed-decisions train split, so this test score is not zero-shot.

Laya base means the upstream convaiinnovations/laya checkpoint. Laya specialist means the upstream convaiinnovations/laya-typed-decisions checkpoint, fine-tuned on the Typed Decisions train split.

The final column uses the pinned test split of Typed Decisions (LocalLLaMA/typed-decisions), with 2,000 decisions and 1,965 untied targets. CalDec Laya used that benchmark's train split; CalDec GLiNER did not. The full comparison is in RESULTS.md.

The test split was used during development. Three exact state values occur in more than one split; the repository's scripts/check_data.py prints the case ids. This limits test independence slightly. The full evaluation contract and per-site table are in RESULTS.md.

Use and limits

The states are synthetic, English, and generally short. Some contain prompt-injection text and invented credentials by design. Treat all state text as untrusted. These are generated cases, not an adversarial safety benchmark. Targets reflect model judgments, including their errors. Evaluate a trained model on real, independently labelled cases before consequential use.

The repository's datasheet documents collection, format, processing, split overlap and intended use. The license is Apache 2.0. Review the hosted model providers' terms when using the recipe to create new data.

Citation and contact

Contact Kristijan Grozdanovski. The repository's CITATION.cff is the machine-readable citation.

@misc{grozdanovski2026assistantdecisions,
  author = {Grozdanovski, Kristijan},
  title = {Assistant Decisions: A Dataset and Recipe for Small Typed-Decision Models},
  year = {2026},
  url = {https://huggingface.co/datasets/kgrozdanovski/assistant-decisions}
}
assistant
calibration
decision-model
synthetic
typed-decisions