MaziyarPanahi/AgentToolDecisions-180K

Dataset

Agent Tool Decisions 180K

10

7 commits

updated Oct 1, 2026

See the code

README

Agent Tool Decisions 180K

180,000 typed decisions that sit behind every agent step. Should the agent call a tool, or answer in text? Which tool? Are the arguments complete? Which tool-using response is better?

An agent spends most of its time making these small decisions, and a frontier LLM is an expensive way to make them. They are typed: a question, a fixed set of choices, one correct answer. That makes them a good fit for small, fast decision models (encoders, classifiers, Jev-style typed decision models, GLiNER2.5-Decide-style heads) that answer in milliseconds on a CPU.

This dataset turns NVIDIA's open agentic data into exactly that format: 180K clean, labeled, split-safe decision rows, ready to train or evaluate a tool router today.

Built on NVIDIA's open data: nvidia/Nemotron-SFT-Agentic-v2 and nvidia/When2Call. Every label comes from the upstream data (programmatic or automated). No output from Jev, TypeSafe or any other proprietary model is included. It contains no clinical or medical data.

Update:

I used HuggingChat+ML Intern to train ModerBERT on this dataset. Everything to run it yourself, prompt, recipe and receipts: https://huggingface.co/MaziyarPanahi/ModernJEV-Decide-Preview

At a glance

Rows180,000
Train / validation / test171,056 / 2,713 / 6,231
Decision groups101,786
Upstream source rows41,809
Leakagenone: no decision group or source row crosses splits
Formatflat Parquet, 35 columns, works with datasets, DuckDB, Polars and pandas

The six decisions

Task familyQuestion the model answersRows
agent_next_action_typeWhat type of action should the assistant take next?79,238
tool_selectionWhich available tool should be called next?37,233
tool_argument_completenessDo the proposed tool calls include every required argument named by the tool schemas?37,233
tool_or_text_actionShould the assistant call a tool or produce a text response next?14,349
tool_response_preferenceWhich response better uses or declines the available tools?8,295
when_to_call_toolWhen2Call-style: should a tool be called at all?3,652

Two answer types (primitive): choice (one label from the criteria, 142,767 rows) and noul (a bounded 0–1 value, 37,233 rows).

Quick start

from datasets import load_dataset

ds = load_dataset("MaziyarPanahi/AgentToolDecisions-180K")
row = ds["train"][0]

print(row["question_text"])   # e.g. "Which available tool should be called next?"
print(row["state_preview"])   # the agent state: conversation + available tools
print(row["criteria_json"])   # the allowed answers
print(row["gold_label"] if row["gold_label"] is not None else row["gold_score"])

Train a tool router on a single task family:

tool_sel = ds["train"].filter(lambda r: r["task_family"] == "tool_selection")
# inputs: question_text + state_json, labels: gold_label (one of the keys in criteria_json)

What you can build with it

  • Tool routers: pick the next tool in milliseconds instead of spending a full LLM call.
  • "Should I call a tool?" gates: cut useless tool calls and the latency they add.
  • Argument checkers: catch missing required arguments before the call fails.
  • Reward and preference models for tool-using responses.
  • Evaluation: measure your agent's decision accuracy per task family, on a leakage-free test split.
  • Race Jev: every row's request_json is a ready-to-send Jev API request ("model": "jev-1.13.0", typed questions with criteria, and the agent state). Run Jev, or any typed decision model, on exactly the same requests and compare against the gold labels.

How to read a row

NeedColumns
Identity and splitrow_id, group_id, output_split
Decision taskpriority, task_family, primitive, question_key
Model inputstate_json, state_preview, question_text, criteria_json
Ready-to-send Jev requestrequest_json
Reference answergold_label, gold_score, gold_json, label_source
Provenancesource_dataset, source_revision, source_row_id, source_license, source_url, transformation, transform_version

Nested payloads are canonical JSON strings, so the Parquet schema stays stable across tasks.

Where the labels come from

Label sourceRows
programmatic_next_message79,238
programmatic_schema_check37,233
programmatic_tool_call37,233
programmatic_message_format14,349
automated_preference_gold8,295
automated_verified_mcq_gold3,652
SourceRows
nvidia/Nemotron-SFT-Agentic-v2153,704
nvidia/When2Call26,296

Validation

An independent audit checks:

  • 180,000 rows with unique row IDs;
  • no decision group or upstream source row crosses splits;
  • valid canonical JSON for state, request, criteria, gold and metadata;
  • choice labels inside the declared criteria, and scores within bounds;
  • a pinned source revision, license, URL and transformation on every row;
  • no proprietary-model prediction, probability, raw-response, usage or latency columns.

Limitations

  • Labels are programmatic or automated upstream labels. They can be noisy or ambiguous, so keep label_source and task_family with every prediction you evaluate.
  • Tool schemas and domains vary across the upstream agent episodes.
  • The typed request format is a normalized transformation, not the original upstream schema.
  • This is a general agent-decision dataset, not a clinical one.

License and attribution

Derived from pinned revisions of nvidia/Nemotron-SFT-Agentic-v2 and nvidia/When2Call. When2Call is marked CC BY 4.0. The Nemotron card lists CC BY 4.0, Apache 2.0 and MIT terms across its components. Every row keeps its recorded source license. Please review and follow the upstream dataset cards when you use or redistribute the data. Thanks to NVIDIA for releasing the source data openly.

agents
decision-models
function-calling
preference-data
routing
structured-prediction
tool-calling
tool-use

MaziyarPanahi/AgentToolDecisions-180K

Dataset

Agent Tool Decisions 180K

10

7 commits

updated Oct 1, 2026

See the code

README

Agent Tool Decisions 180K

180,000 typed decisions that sit behind every agent step. Should the agent call a tool, or answer in text? Which tool? Are the arguments complete? Which tool-using response is better?

An agent spends most of its time making these small decisions, and a frontier LLM is an expensive way to make them. They are typed: a question, a fixed set of choices, one correct answer. That makes them a good fit for small, fast decision models (encoders, classifiers, Jev-style typed decision models, GLiNER2.5-Decide-style heads) that answer in milliseconds on a CPU.

This dataset turns NVIDIA's open agentic data into exactly that format: 180K clean, labeled, split-safe decision rows, ready to train or evaluate a tool router today.

Built on NVIDIA's open data: nvidia/Nemotron-SFT-Agentic-v2 and nvidia/When2Call. Every label comes from the upstream data (programmatic or automated). No output from Jev, TypeSafe or any other proprietary model is included. It contains no clinical or medical data.

Update:

I used HuggingChat+ML Intern to train ModerBERT on this dataset. Everything to run it yourself, prompt, recipe and receipts: https://huggingface.co/MaziyarPanahi/ModernJEV-Decide-Preview

At a glance

Rows180,000
Train / validation / test171,056 / 2,713 / 6,231
Decision groups101,786
Upstream source rows41,809
Leakagenone: no decision group or source row crosses splits
Formatflat Parquet, 35 columns, works with datasets, DuckDB, Polars and pandas

The six decisions

Task familyQuestion the model answersRows
agent_next_action_typeWhat type of action should the assistant take next?79,238
tool_selectionWhich available tool should be called next?37,233
tool_argument_completenessDo the proposed tool calls include every required argument named by the tool schemas?37,233
tool_or_text_actionShould the assistant call a tool or produce a text response next?14,349
tool_response_preferenceWhich response better uses or declines the available tools?8,295
when_to_call_toolWhen2Call-style: should a tool be called at all?3,652

Two answer types (primitive): choice (one label from the criteria, 142,767 rows) and noul (a bounded 0–1 value, 37,233 rows).

Quick start

from datasets import load_dataset

ds = load_dataset("MaziyarPanahi/AgentToolDecisions-180K")
row = ds["train"][0]

print(row["question_text"])   # e.g. "Which available tool should be called next?"
print(row["state_preview"])   # the agent state: conversation + available tools
print(row["criteria_json"])   # the allowed answers
print(row["gold_label"] if row["gold_label"] is not None else row["gold_score"])

Train a tool router on a single task family:

tool_sel = ds["train"].filter(lambda r: r["task_family"] == "tool_selection")
# inputs: question_text + state_json, labels: gold_label (one of the keys in criteria_json)

What you can build with it

  • Tool routers: pick the next tool in milliseconds instead of spending a full LLM call.
  • "Should I call a tool?" gates: cut useless tool calls and the latency they add.
  • Argument checkers: catch missing required arguments before the call fails.
  • Reward and preference models for tool-using responses.
  • Evaluation: measure your agent's decision accuracy per task family, on a leakage-free test split.
  • Race Jev: every row's request_json is a ready-to-send Jev API request ("model": "jev-1.13.0", typed questions with criteria, and the agent state). Run Jev, or any typed decision model, on exactly the same requests and compare against the gold labels.

How to read a row

NeedColumns
Identity and splitrow_id, group_id, output_split
Decision taskpriority, task_family, primitive, question_key
Model inputstate_json, state_preview, question_text, criteria_json
Ready-to-send Jev requestrequest_json
Reference answergold_label, gold_score, gold_json, label_source
Provenancesource_dataset, source_revision, source_row_id, source_license, source_url, transformation, transform_version

Nested payloads are canonical JSON strings, so the Parquet schema stays stable across tasks.

Where the labels come from

Label sourceRows
programmatic_next_message79,238
programmatic_schema_check37,233
programmatic_tool_call37,233
programmatic_message_format14,349
automated_preference_gold8,295
automated_verified_mcq_gold3,652
SourceRows
nvidia/Nemotron-SFT-Agentic-v2153,704
nvidia/When2Call26,296

Validation

An independent audit checks:

  • 180,000 rows with unique row IDs;
  • no decision group or upstream source row crosses splits;
  • valid canonical JSON for state, request, criteria, gold and metadata;
  • choice labels inside the declared criteria, and scores within bounds;
  • a pinned source revision, license, URL and transformation on every row;
  • no proprietary-model prediction, probability, raw-response, usage or latency columns.

Limitations

  • Labels are programmatic or automated upstream labels. They can be noisy or ambiguous, so keep label_source and task_family with every prediction you evaluate.
  • Tool schemas and domains vary across the upstream agent episodes.
  • The typed request format is a normalized transformation, not the original upstream schema.
  • This is a general agent-decision dataset, not a clinical one.

License and attribution

Derived from pinned revisions of nvidia/Nemotron-SFT-Agentic-v2 and nvidia/When2Call. When2Call is marked CC BY 4.0. The Nemotron card lists CC BY 4.0, Apache 2.0 and MIT terms across its components. Every row keeps its recorded source license. Please review and follow the upstream dataset cards when you use or redistribute the data. Thanks to NVIDIA for releasing the source data openly.

agents
decision-models
function-calling
preference-data
routing
structured-prediction
tool-calling
tool-use