sappho192/aihub-translation-dataset

Code to merge some datasets in AIHub into single dataset

Python

2

15 commits

updated Nov 2, 2025

See the code

README

AIHub Translation Dataset Merger

A Python tool for merging multiple AIHub translation datasets while preserving the original language direction. The tool processes various AIHub translation corpus datasets and outputs them grouped by language pair (source→target) in multiple formats.

Features

  • Language Direction Preservation: Maintains original source→target language pairs (ko↔ja, ko→en)
  • Multiple Language Support: Processes Korean-Japanese and Korean-English translations
  • Multiple Output Formats: Generates JSON, JSONL, and CSV files
  • Organized Output: Datasets are saved in separate directories by language pair
  • Multiple Dataset Support: Processes 5 different AIHub translation datasets
  • Train/Validation Splits: Preserves original train and validation splits

Supported Datasets

The tool currently supports these AIHub datasets:

  1. Dataset 027: 일상생활 및 구어체 한-중, 한-일 번역 병렬 말뭉치 데이터
    • Provides both ja→ko and ko→ja translations
  2. Dataset 053: 한국어-다국어 번역 말뭉치(기술과학)
    • Provides ko→ja translations
  3. Dataset 054: 한국어-다국어 번역 말뭉치(기초과학)
    • Provides ko→ja and ko→en translations
  4. Dataset 055: 한국어-다국어 번역 말뭉치(인문학)
    • Provides ko→ja and ko→en translations
  5. 한국어-일본어 번역 말뭉치
    • Provides ko→ja translations

Installation

This project uses uv for dependency management:

# Install uv if you haven't already
pip install uv

# Clone the repository
git clone <repository-url>
cd aihub-translation-dataset

Usage

Run the main script with the path to your AIHub datasets:

uv run main.py --dataset-path /path/to/aihub

On Windows:

uv run main.py --dataset-path K:\DATASET\aihub

To specify a custom output directory:

uv run main.py --dataset-path /path/to/aihub --output-path /path/to/output

Command-line Arguments

  • --dataset-path (required): Root path to the AIHub datasets directory
  • --output-path (optional): Output path for merged datasets. Defaults to <dataset-path>/merged if not specified

Required Directory Structure

Your AIHub root directory should contain:

/path/to/aihub/
  ├── 027.일상생활 및 구어체 한-중, 한-일 번역 병렬 말뭉치 데이터/
  ├── 053.한국어-다국어(영어 제외) 번역 말뭉치(기술과학)/
  ├── 054.한국어-다국어 번역 말뭉치(기초과학)/
  ├── 055.한국어-다국어 번역 말뭉치(인문학)/
  └── 한국어-일본어 번역 말뭉치/

Output Structure

The merged datasets will be saved to the specified output directory (default: output/) organized by language pair:

<output-path>/
  ├── ko_ja/              # Korean→Japanese translations
  │   ├── train.json
  │   ├── train.jsonl
  │   ├── train.csv
  │   ├── val.json
  │   ├── val.jsonl
  │   └── val.csv
  ├── ja_ko/              # Japanese→Korean translations
  │   ├── train.json
  │   ├── train.jsonl
  │   ├── train.csv
  │   ├── val.json
  │   ├── val.jsonl
  │   └── val.csv
  └── ko_en/              # Korean→English translations
      ├── train.json
      ├── train.jsonl
      ├── train.csv
      ├── val.json
      ├── val.jsonl
      └── val.csv

Output Formats

All formats contain the same data with sourceString and targetString fields:

  • JSON: Array of translation objects with indentation
  • JSONL: One JSON object per line (JSON Lines format)
  • CSV: CSV with header row (sourceString,targetString)

Data Format

Each translation pair is saved as:

{
  "sourceString": "source text in source language",
  "targetString": "target text in target language"
}

Architecture

The codebase is organized as follows:

  • AiHub/AiHubDatasetBase.py: Base class for all dataset implementations
  • AiHub/AiHubDataset*.py: Dataset-specific implementations
  • AiHub/util/DatasetGenerator.py: Utilities for merging and saving datasets
  • main.py: Main script that orchestrates the merging process

Adding New Datasets

To add support for a new AIHub dataset:

  1. Create a new class in AiHub/ inheriting from AiHubDatasetBase
  2. Implement the required methods (make_train_dataset(), make_val_dataset())
  3. Return data in the language pair format with correct source_lang and target_lang
  4. Add the dataset to main.py

See CLAUDE.md for detailed architecture documentation.

License

MIT License

sappho192/aihub-translation-dataset

Code to merge some datasets in AIHub into single dataset

Python

2

15 commits

updated Nov 2, 2025

See the code

README

AIHub Translation Dataset Merger

A Python tool for merging multiple AIHub translation datasets while preserving the original language direction. The tool processes various AIHub translation corpus datasets and outputs them grouped by language pair (source→target) in multiple formats.

Features

  • Language Direction Preservation: Maintains original source→target language pairs (ko↔ja, ko→en)
  • Multiple Language Support: Processes Korean-Japanese and Korean-English translations
  • Multiple Output Formats: Generates JSON, JSONL, and CSV files
  • Organized Output: Datasets are saved in separate directories by language pair
  • Multiple Dataset Support: Processes 5 different AIHub translation datasets
  • Train/Validation Splits: Preserves original train and validation splits

Supported Datasets

The tool currently supports these AIHub datasets:

  1. Dataset 027: 일상생활 및 구어체 한-중, 한-일 번역 병렬 말뭉치 데이터
    • Provides both ja→ko and ko→ja translations
  2. Dataset 053: 한국어-다국어 번역 말뭉치(기술과학)
    • Provides ko→ja translations
  3. Dataset 054: 한국어-다국어 번역 말뭉치(기초과학)
    • Provides ko→ja and ko→en translations
  4. Dataset 055: 한국어-다국어 번역 말뭉치(인문학)
    • Provides ko→ja and ko→en translations
  5. 한국어-일본어 번역 말뭉치
    • Provides ko→ja translations

Installation

This project uses uv for dependency management:

# Install uv if you haven't already
pip install uv

# Clone the repository
git clone <repository-url>
cd aihub-translation-dataset

Usage

Run the main script with the path to your AIHub datasets:

uv run main.py --dataset-path /path/to/aihub

On Windows:

uv run main.py --dataset-path K:\DATASET\aihub

To specify a custom output directory:

uv run main.py --dataset-path /path/to/aihub --output-path /path/to/output

Command-line Arguments

  • --dataset-path (required): Root path to the AIHub datasets directory
  • --output-path (optional): Output path for merged datasets. Defaults to <dataset-path>/merged if not specified

Required Directory Structure

Your AIHub root directory should contain:

/path/to/aihub/
  ├── 027.일상생활 및 구어체 한-중, 한-일 번역 병렬 말뭉치 데이터/
  ├── 053.한국어-다국어(영어 제외) 번역 말뭉치(기술과학)/
  ├── 054.한국어-다국어 번역 말뭉치(기초과학)/
  ├── 055.한국어-다국어 번역 말뭉치(인문학)/
  └── 한국어-일본어 번역 말뭉치/

Output Structure

The merged datasets will be saved to the specified output directory (default: output/) organized by language pair:

<output-path>/
  ├── ko_ja/              # Korean→Japanese translations
  │   ├── train.json
  │   ├── train.jsonl
  │   ├── train.csv
  │   ├── val.json
  │   ├── val.jsonl
  │   └── val.csv
  ├── ja_ko/              # Japanese→Korean translations
  │   ├── train.json
  │   ├── train.jsonl
  │   ├── train.csv
  │   ├── val.json
  │   ├── val.jsonl
  │   └── val.csv
  └── ko_en/              # Korean→English translations
      ├── train.json
      ├── train.jsonl
      ├── train.csv
      ├── val.json
      ├── val.jsonl
      └── val.csv

Output Formats

All formats contain the same data with sourceString and targetString fields:

  • JSON: Array of translation objects with indentation
  • JSONL: One JSON object per line (JSON Lines format)
  • CSV: CSV with header row (sourceString,targetString)

Data Format

Each translation pair is saved as:

{
  "sourceString": "source text in source language",
  "targetString": "target text in target language"
}

Architecture

The codebase is organized as follows:

  • AiHub/AiHubDatasetBase.py: Base class for all dataset implementations
  • AiHub/AiHubDataset*.py: Dataset-specific implementations
  • AiHub/util/DatasetGenerator.py: Utilities for merging and saving datasets
  • main.py: Main script that orchestrates the merging process

Adding New Datasets

To add support for a new AIHub dataset:

  1. Create a new class in AiHub/ inheriting from AiHubDatasetBase
  2. Implement the required methods (make_train_dataset(), make_val_dataset())
  3. Return data in the language pair format with correct source_lang and target_lang
  4. Add the dataset to main.py

See CLAUDE.md for detailed architecture documentation.

License

MIT License

Languages

Python

100.0%