Time-HD-Anonymous/ST-Bench

Dataset

ST-Bench: Spatial-Temporal Reasoning Benchmark

1

6 commits

1 linked in READMEs

updated Jan 6, 2026

See the code

README

ST-Bench: Spatial-Temporal Reasoning Benchmark

ST-Bench is a comprehensive benchmark dataset for training and evaluating spatial-temporal reasoning capabilities in large language models. It includes data with raw time series, text descriptions, and image visualizations.

πŸ“Š Dataset Overview

Default Data (with time_series key)

SubsetDescriptionFilesTotal Size
ST-AlignAlignment data for initial training3 files~3.2GB
ST-CausalCausal reasoning data1 file~14MB
ST-SFTSupervised fine-tuning data4 files~50MB
ST-CoTChain-of-thought reasoning data4 files~191MB
ST-RLReinforcement learning data4 files~50MB
ST-TestTest/evaluation data4 files~24MB

Text-based Data (text description only)

SubsetDescriptionFilesTotal Size
ST-SFT-TextSupervised fine-tuning data4 files~50MB
ST-CoT-TextChain-of-thought reasoning data4 files~194MB
ST-RL-TextReinforcement learning data4 files~52MB
ST-Test-TextTest/evaluation data4 files~24MB

Image-based Data (with image reference)

SubsetDescriptionFilesTotal Size
ST-CoT-ImageChain-of-thought with visualizations4 files~61MB
ST-RL-ImageRL data with visualizations4 files~6.7MB

Note: The visualization images referenced in Image-based data are not included in this release due to size constraints.

πŸ“ Directory Structure

ST-Bench/
β”œβ”€β”€ ST-Align/                    # Alignment data
β”‚   β”œβ”€β”€ alignment.jsonl
β”‚   β”œβ”€β”€ alignment_train.jsonl
β”‚   └── alignment_test.jsonl
β”‚
β”œβ”€β”€ ST-Causal/                   # Causal reasoning
β”‚   └── causal.jsonl
β”‚
β”œβ”€β”€ ST-SFT/                      # Supervised Fine-Tuning (Default)
β”œβ”€β”€ ST-CoT/                      # Chain-of-Thought (Default)
β”œβ”€β”€ ST-RL/                       # Reinforcement Learning (Default)
β”œβ”€β”€ ST-Test/                     # Test Set (Default)
β”‚
β”œβ”€β”€ ST-SFT-Text/                 # Supervised Fine-Tuning (Text)
β”œβ”€β”€ ST-CoT-Text/                 # Chain-of-Thought (Text)
β”œβ”€β”€ ST-RL-Text/                  # Reinforcement Learning (Text)
β”œβ”€β”€ ST-Test-Text/                # Test Set (Text)
β”‚
β”œβ”€β”€ ST-CoT-Image/                # Chain-of-Thought (Image)
└── ST-RL-Image/                 # Reinforcement Learning (Image)

Each reasoning folder contains:

  • correlation_*.jsonl - Correlation reasoning tasks
  • entity_*.jsonl - Entity reasoning tasks
  • etiological_*.jsonl - Etiological (cause-effect) reasoning tasks
  • forecasting_*.jsonl - Forecasting tasks

🎯 Task Types

The reasoning tasks cover four main categories:

  1. Correlation: Understanding relationships between time series variables
  2. Entity: Identifying and reasoning about entities in temporal data
  3. Etiological: Cause-effect reasoning in temporal sequences
  4. Forecasting: Predicting future values based on historical patterns

πŸ“– Usage

from datasets import load_dataset

# Load default data (with time_series key)
cot_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT")

# Load text-only data
cot_text = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Text")

# Load image-based data
cot_image = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Image")

# Load alignment data
align_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-Align")

πŸ“œ License

This dataset is released under the Apache 2.0 License.

πŸ“ Citation

If you use this dataset in your research, please cite:

@misc{stbench2025,
  title={ST-Bench: A Benchmark for Spatial-Temporal Reasoning},
  author={Anonymous},
  year={2025},
  publisher={Hugging Face}
}
multimodal
spatio-temporal-reasoning
time-series

Time-HD-Anonymous/ST-Bench

Dataset

ST-Bench: Spatial-Temporal Reasoning Benchmark

1

6 commits

1 linked in READMEs

updated Jan 6, 2026

See the code

README

ST-Bench: Spatial-Temporal Reasoning Benchmark

ST-Bench is a comprehensive benchmark dataset for training and evaluating spatial-temporal reasoning capabilities in large language models. It includes data with raw time series, text descriptions, and image visualizations.

πŸ“Š Dataset Overview

Default Data (with time_series key)

SubsetDescriptionFilesTotal Size
ST-AlignAlignment data for initial training3 files~3.2GB
ST-CausalCausal reasoning data1 file~14MB
ST-SFTSupervised fine-tuning data4 files~50MB
ST-CoTChain-of-thought reasoning data4 files~191MB
ST-RLReinforcement learning data4 files~50MB
ST-TestTest/evaluation data4 files~24MB

Text-based Data (text description only)

SubsetDescriptionFilesTotal Size
ST-SFT-TextSupervised fine-tuning data4 files~50MB
ST-CoT-TextChain-of-thought reasoning data4 files~194MB
ST-RL-TextReinforcement learning data4 files~52MB
ST-Test-TextTest/evaluation data4 files~24MB

Image-based Data (with image reference)

SubsetDescriptionFilesTotal Size
ST-CoT-ImageChain-of-thought with visualizations4 files~61MB
ST-RL-ImageRL data with visualizations4 files~6.7MB

Note: The visualization images referenced in Image-based data are not included in this release due to size constraints.

πŸ“ Directory Structure

ST-Bench/
β”œβ”€β”€ ST-Align/                    # Alignment data
β”‚   β”œβ”€β”€ alignment.jsonl
β”‚   β”œβ”€β”€ alignment_train.jsonl
β”‚   └── alignment_test.jsonl
β”‚
β”œβ”€β”€ ST-Causal/                   # Causal reasoning
β”‚   └── causal.jsonl
β”‚
β”œβ”€β”€ ST-SFT/                      # Supervised Fine-Tuning (Default)
β”œβ”€β”€ ST-CoT/                      # Chain-of-Thought (Default)
β”œβ”€β”€ ST-RL/                       # Reinforcement Learning (Default)
β”œβ”€β”€ ST-Test/                     # Test Set (Default)
β”‚
β”œβ”€β”€ ST-SFT-Text/                 # Supervised Fine-Tuning (Text)
β”œβ”€β”€ ST-CoT-Text/                 # Chain-of-Thought (Text)
β”œβ”€β”€ ST-RL-Text/                  # Reinforcement Learning (Text)
β”œβ”€β”€ ST-Test-Text/                # Test Set (Text)
β”‚
β”œβ”€β”€ ST-CoT-Image/                # Chain-of-Thought (Image)
└── ST-RL-Image/                 # Reinforcement Learning (Image)

Each reasoning folder contains:

  • correlation_*.jsonl - Correlation reasoning tasks
  • entity_*.jsonl - Entity reasoning tasks
  • etiological_*.jsonl - Etiological (cause-effect) reasoning tasks
  • forecasting_*.jsonl - Forecasting tasks

🎯 Task Types

The reasoning tasks cover four main categories:

  1. Correlation: Understanding relationships between time series variables
  2. Entity: Identifying and reasoning about entities in temporal data
  3. Etiological: Cause-effect reasoning in temporal sequences
  4. Forecasting: Predicting future values based on historical patterns

πŸ“– Usage

from datasets import load_dataset

# Load default data (with time_series key)
cot_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT")

# Load text-only data
cot_text = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Text")

# Load image-based data
cot_image = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-CoT-Image")

# Load alignment data
align_data = load_dataset("Time-HD-Anonymous/ST-Bench", "ST-Align")

πŸ“œ License

This dataset is released under the Apache 2.0 License.

πŸ“ Citation

If you use this dataset in your research, please cite:

@misc{stbench2025,
  title={ST-Bench: A Benchmark for Spatial-Temporal Reasoning},
  author={Anonymous},
  year={2025},
  publisher={Hugging Face}
}
multimodal
spatio-temporal-reasoning
time-series