180,000 typed decisions that sit behind every agent step. Should the agent call a tool, or answer in text? Which tool? Are the arguments complete? Which tool-using response is better?
An agent spends most of its time making these small decisions, and a frontier LLM is an expensive way to make them. They are typed: a question, a fixed set of choices, one correct answer. That makes them a good fit for small, fast decision models (encoders, classifiers, Jev-style typed decision models, GLiNER2.5-Decide-style heads) that answer in milliseconds on a CPU.
This dataset turns NVIDIA's open agentic data into exactly that format: 180K clean, labeled, split-safe decision rows, ready to train or evaluate a tool router today.
Built on NVIDIA's open data:
nvidia/Nemotron-SFT-Agentic-v2andnvidia/When2Call. Every label comes from the upstream data (programmatic or automated). No output from Jev, TypeSafe or any other proprietary model is included. It contains no clinical or medical data.
I used HuggingChat+ML Intern to train ModerBERT on this dataset. Everything to run it yourself, prompt, recipe and receipts: https://huggingface.co/MaziyarPanahi/ModernJEV-Decide-Preview
| Rows | 180,000 |
| Train / validation / test | 171,056 / 2,713 / 6,231 |
| Decision groups | 101,786 |
| Upstream source rows | 41,809 |
| Leakage | none: no decision group or source row crosses splits |
| Format | flat Parquet, 35 columns, works with datasets, DuckDB, Polars and pandas |
| Task family | Question the model answers | Rows |
|---|---|---|
agent_next_action_type | What type of action should the assistant take next? | 79,238 |
tool_selection | Which available tool should be called next? | 37,233 |
tool_argument_completeness | Do the proposed tool calls include every required argument named by the tool schemas? | 37,233 |
tool_or_text_action | Should the assistant call a tool or produce a text response next? | 14,349 |
tool_response_preference | Which response better uses or declines the available tools? | 8,295 |
when_to_call_tool | When2Call-style: should a tool be called at all? | 3,652 |
Two answer types (primitive): choice (one label from the criteria, 142,767 rows) and
noul (a bounded 0–1 value, 37,233 rows).
from datasets import load_dataset
ds = load_dataset("MaziyarPanahi/AgentToolDecisions-180K")
row = ds["train"][0]
print(row["question_text"]) # e.g. "Which available tool should be called next?"
print(row["state_preview"]) # the agent state: conversation + available tools
print(row["criteria_json"]) # the allowed answers
print(row["gold_label"] if row["gold_label"] is not None else row["gold_score"])
Train a tool router on a single task family:
tool_sel = ds["train"].filter(lambda r: r["task_family"] == "tool_selection")
# inputs: question_text + state_json, labels: gold_label (one of the keys in criteria_json)
request_json is a ready-to-send Jev API request ("model": "jev-1.13.0",
typed questions with criteria, and the agent state). Run Jev, or any typed decision model, on
exactly the same requests and compare against the gold labels.| Need | Columns |
|---|---|
| Identity and split | row_id, group_id, output_split |
| Decision task | priority, task_family, primitive, question_key |
| Model input | state_json, state_preview, question_text, criteria_json |
| Ready-to-send Jev request | request_json |
| Reference answer | gold_label, gold_score, gold_json, label_source |
| Provenance | source_dataset, source_revision, source_row_id, source_license, source_url, transformation, transform_version |
Nested payloads are canonical JSON strings, so the Parquet schema stays stable across tasks.
| Label source | Rows |
|---|---|
programmatic_next_message | 79,238 |
programmatic_schema_check | 37,233 |
programmatic_tool_call | 37,233 |
programmatic_message_format | 14,349 |
automated_preference_gold | 8,295 |
automated_verified_mcq_gold | 3,652 |
| Source | Rows |
|---|---|
nvidia/Nemotron-SFT-Agentic-v2 | 153,704 |
nvidia/When2Call | 26,296 |
An independent audit checks:
label_source and task_family with every prediction you evaluate.Derived from pinned revisions of nvidia/Nemotron-SFT-Agentic-v2 and nvidia/When2Call. When2Call
is marked CC BY 4.0. The Nemotron card lists CC BY 4.0, Apache 2.0 and MIT terms across its
components. Every row keeps its recorded source license. Please review and follow the upstream
dataset cards when you use or redistribute the data. Thanks to NVIDIA for releasing the source data
openly.
180,000 typed decisions that sit behind every agent step. Should the agent call a tool, or answer in text? Which tool? Are the arguments complete? Which tool-using response is better?
An agent spends most of its time making these small decisions, and a frontier LLM is an expensive way to make them. They are typed: a question, a fixed set of choices, one correct answer. That makes them a good fit for small, fast decision models (encoders, classifiers, Jev-style typed decision models, GLiNER2.5-Decide-style heads) that answer in milliseconds on a CPU.
This dataset turns NVIDIA's open agentic data into exactly that format: 180K clean, labeled, split-safe decision rows, ready to train or evaluate a tool router today.
Built on NVIDIA's open data:
nvidia/Nemotron-SFT-Agentic-v2andnvidia/When2Call. Every label comes from the upstream data (programmatic or automated). No output from Jev, TypeSafe or any other proprietary model is included. It contains no clinical or medical data.
I used HuggingChat+ML Intern to train ModerBERT on this dataset. Everything to run it yourself, prompt, recipe and receipts: https://huggingface.co/MaziyarPanahi/ModernJEV-Decide-Preview
| Rows | 180,000 |
| Train / validation / test | 171,056 / 2,713 / 6,231 |
| Decision groups | 101,786 |
| Upstream source rows | 41,809 |
| Leakage | none: no decision group or source row crosses splits |
| Format | flat Parquet, 35 columns, works with datasets, DuckDB, Polars and pandas |
| Task family | Question the model answers | Rows |
|---|---|---|
agent_next_action_type | What type of action should the assistant take next? | 79,238 |
tool_selection | Which available tool should be called next? | 37,233 |
tool_argument_completeness | Do the proposed tool calls include every required argument named by the tool schemas? | 37,233 |
tool_or_text_action | Should the assistant call a tool or produce a text response next? | 14,349 |
tool_response_preference | Which response better uses or declines the available tools? | 8,295 |
when_to_call_tool | When2Call-style: should a tool be called at all? | 3,652 |
Two answer types (primitive): choice (one label from the criteria, 142,767 rows) and
noul (a bounded 0–1 value, 37,233 rows).
from datasets import load_dataset
ds = load_dataset("MaziyarPanahi/AgentToolDecisions-180K")
row = ds["train"][0]
print(row["question_text"]) # e.g. "Which available tool should be called next?"
print(row["state_preview"]) # the agent state: conversation + available tools
print(row["criteria_json"]) # the allowed answers
print(row["gold_label"] if row["gold_label"] is not None else row["gold_score"])
Train a tool router on a single task family:
tool_sel = ds["train"].filter(lambda r: r["task_family"] == "tool_selection")
# inputs: question_text + state_json, labels: gold_label (one of the keys in criteria_json)
request_json is a ready-to-send Jev API request ("model": "jev-1.13.0",
typed questions with criteria, and the agent state). Run Jev, or any typed decision model, on
exactly the same requests and compare against the gold labels.| Need | Columns |
|---|---|
| Identity and split | row_id, group_id, output_split |
| Decision task | priority, task_family, primitive, question_key |
| Model input | state_json, state_preview, question_text, criteria_json |
| Ready-to-send Jev request | request_json |
| Reference answer | gold_label, gold_score, gold_json, label_source |
| Provenance | source_dataset, source_revision, source_row_id, source_license, source_url, transformation, transform_version |
Nested payloads are canonical JSON strings, so the Parquet schema stays stable across tasks.
| Label source | Rows |
|---|---|
programmatic_next_message | 79,238 |
programmatic_schema_check | 37,233 |
programmatic_tool_call | 37,233 |
programmatic_message_format | 14,349 |
automated_preference_gold | 8,295 |
automated_verified_mcq_gold | 3,652 |
| Source | Rows |
|---|---|
nvidia/Nemotron-SFT-Agentic-v2 | 153,704 |
nvidia/When2Call | 26,296 |
An independent audit checks:
label_source and task_family with every prediction you evaluate.Derived from pinned revisions of nvidia/Nemotron-SFT-Agentic-v2 and nvidia/When2Call. When2Call
is marked CC BY 4.0. The Nemotron card lists CC BY 4.0, Apache 2.0 and MIT terms across its
components. Every row keeps its recorded source license. Please review and follow the upstream
dataset cards when you use or redistribute the data. Thanks to NVIDIA for releasing the source data
openly.