ST-Bench: Spatial-Temporal Reasoning Benchmark
1
6 commits
1 linked in READMEs
updated Jan 6, 2026
ST-Bench is a comprehensive benchmark dataset for training and evaluating spatial-temporal reasoning capabilities in large language models. It includes data with raw time series, text descriptions, and image visualizations.
time_series key)| Subset | Description | Files | Total Size |
|---|---|---|---|
| ST-Align | Alignment data for initial training | 3 files | ~3.2GB |
| ST-Causal | Causal reasoning data | 1 file | ~14MB |
| ST-SFT | Supervised fine-tuning data | 4 files | ~50MB |
| ST-CoT | Chain-of-thought reasoning data | 4 files | ~191MB |
| ST-RL | Reinforcement learning data | 4 files | ~50MB |
| ST-Test | Test/evaluation data | 4 files | ~24MB |
| Subset | Description | Files | Total Size |
|---|---|---|---|
| ST-SFT-Text | Supervised fine-tuning data | 4 files | ~50MB |
| ST-CoT-Text | Chain-of-thought reasoning data | 4 files | ~194MB |
| ST-RL-Text | Reinforcement learning data | 4 files | ~52MB |
| ST-Test-Text | Test/evaluation data | 4 files | ~24MB |
| Subset | Description | Files | Total Size |
|---|---|---|---|
| ST-CoT-Image | Chain-of-thought with visualizations | 4 files | ~61MB |
| ST-RL-Image | RL data with visualizations | 4 files | ~6.7MB |
Note: The visualization images referenced in Image-based data are not included in this release due to size constraints.
ST-Bench/
βββ ST-Align/ # Alignment data
β βββ alignment.jsonl
β βββ alignment_train.jsonl
β βββ alignment_test.jsonl
β
βββ ST-Causal/ # Causal reasoning
β βββ causal.jsonl
β
βββ ST-SFT/ # Supervised Fine-Tuning (Default)
βββ ST-CoT/ # Chain-of-Thought (Default)
βββ ST-RL/ # Reinforcement Learning (Default)
βββ ST-Test/ # Test Set (Default)
β
βββ ST-SFT-Text/ # Supervised Fine-Tuning (Text)
βββ ST-CoT-Text/ # Chain-of-Thought (Text)
βββ ST-RL-Text/ # Reinforcement Learning (Text)
βββ ST-Test-Text/ # Test Set (Text)
β
βββ ST-CoT-Image/ # Chain-of-Thought (Image)
βββ ST-RL-Image/ # Reinforcement Learning (Image)
Each reasoning folder contains:
correlation_*.jsonl - Correlation reasoning tasksentity_*.jsonl - Entity reasoning tasksetiological_*.jsonl - Etiological (cause-effect) reasoning tasksforecasting_*.jsonl - Forecasting tasksThe reasoning tasks cover four main categories:
from datasets import load_dataset
# Load default data (with time_series key)
cot_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT")
# Load text-only data
cot_text = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Text")
# Load image-based data
cot_image = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Image")
# Load alignment data
align_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-Align")
This dataset is released under the Apache 2.0 License.
If you use this dataset in your research, please cite:
@misc{stbench2025,
title={ST-Bench: A Benchmark for Spatial-Temporal Reasoning},
author={Anonymous},
year={2025},
publisher={Hugging Face}
}
ST-Bench: Spatial-Temporal Reasoning Benchmark
1
6 commits
1 linked in READMEs
updated Jan 6, 2026
ST-Bench is a comprehensive benchmark dataset for training and evaluating spatial-temporal reasoning capabilities in large language models. It includes data with raw time series, text descriptions, and image visualizations.
time_series key)| Subset | Description | Files | Total Size |
|---|---|---|---|
| ST-Align | Alignment data for initial training | 3 files | ~3.2GB |
| ST-Causal | Causal reasoning data | 1 file | ~14MB |
| ST-SFT | Supervised fine-tuning data | 4 files | ~50MB |
| ST-CoT | Chain-of-thought reasoning data | 4 files | ~191MB |
| ST-RL | Reinforcement learning data | 4 files | ~50MB |
| ST-Test | Test/evaluation data | 4 files | ~24MB |
| Subset | Description | Files | Total Size |
|---|---|---|---|
| ST-SFT-Text | Supervised fine-tuning data | 4 files | ~50MB |
| ST-CoT-Text | Chain-of-thought reasoning data | 4 files | ~194MB |
| ST-RL-Text | Reinforcement learning data | 4 files | ~52MB |
| ST-Test-Text | Test/evaluation data | 4 files | ~24MB |
| Subset | Description | Files | Total Size |
|---|---|---|---|
| ST-CoT-Image | Chain-of-thought with visualizations | 4 files | ~61MB |
| ST-RL-Image | RL data with visualizations | 4 files | ~6.7MB |
Note: The visualization images referenced in Image-based data are not included in this release due to size constraints.
ST-Bench/
βββ ST-Align/ # Alignment data
β βββ alignment.jsonl
β βββ alignment_train.jsonl
β βββ alignment_test.jsonl
β
βββ ST-Causal/ # Causal reasoning
β βββ causal.jsonl
β
βββ ST-SFT/ # Supervised Fine-Tuning (Default)
βββ ST-CoT/ # Chain-of-Thought (Default)
βββ ST-RL/ # Reinforcement Learning (Default)
βββ ST-Test/ # Test Set (Default)
β
βββ ST-SFT-Text/ # Supervised Fine-Tuning (Text)
βββ ST-CoT-Text/ # Chain-of-Thought (Text)
βββ ST-RL-Text/ # Reinforcement Learning (Text)
βββ ST-Test-Text/ # Test Set (Text)
β
βββ ST-CoT-Image/ # Chain-of-Thought (Image)
βββ ST-RL-Image/ # Reinforcement Learning (Image)
Each reasoning folder contains:
correlation_*.jsonl - Correlation reasoning tasksentity_*.jsonl - Entity reasoning tasksetiological_*.jsonl - Etiological (cause-effect) reasoning tasksforecasting_*.jsonl - Forecasting tasksThe reasoning tasks cover four main categories:
from datasets import load_dataset
# Load default data (with time_series key)
cot_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT")
# Load text-only data
cot_text = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Text")
# Load image-based data
cot_image = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Image")
# Load alignment data
align_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-Align")
This dataset is released under the Apache 2.0 License.
If you use this dataset in your research, please cite:
@misc{stbench2025,
title={ST-Bench: A Benchmark for Spatial-Temporal Reasoning},
author={Anonymous},
year={2025},
publisher={Hugging Face}
}