This dataset is derived from the work presented in SCRIBE: Structured Chain Reasoning for Interactive Behavior Explanations using Tool Calling (2025). It contains training and evaluation data for developing and benchmarking multi-hop, tool-augmented reasoning models in educational settings.
SCRIBE introduces a framework where smaller open-source LLMs are fine-tuned to provide pedagogically valid, personalized student feedback through iterative reasoning and tool calls. The dataset supports training such models through synthetic but realistic student–feedback interactions.
We provide four splits, reflecting two stages of fine-tuning and two distinct evaluation sets:
Each example includes:
Courses included:
4 commits
This dataset is derived from the work presented in SCRIBE: Structured Chain Reasoning for Interactive Behavior Explanations using Tool Calling (2025). It contains training and evaluation data for developing and benchmarking multi-hop, tool-augmented reasoning models in educational settings.
SCRIBE introduces a framework where smaller open-source LLMs are fine-tuned to provide pedagogically valid, personalized student feedback through iterative reasoning and tool calls. The dataset supports training such models through synthetic but realistic student–feedback interactions.
We provide four splits, reflecting two stages of fine-tuning and two distinct evaluation sets:
Each example includes:
Courses included:
4 commits