Code to merge some datasets in AIHub into single dataset
Python
2
15 commits
updated Nov 2, 2025
A Python tool for merging multiple AIHub translation datasets while preserving the original language direction. The tool processes various AIHub translation corpus datasets and outputs them grouped by language pair (source→target) in multiple formats.
The tool currently supports these AIHub datasets:
This project uses uv for dependency management:
# Install uv if you haven't already
pip install uv
# Clone the repository
git clone <repository-url>
cd aihub-translation-dataset
Run the main script with the path to your AIHub datasets:
uv run main.py --dataset-path /path/to/aihub
On Windows:
uv run main.py --dataset-path K:\DATASET\aihub
To specify a custom output directory:
uv run main.py --dataset-path /path/to/aihub --output-path /path/to/output
--dataset-path (required): Root path to the AIHub datasets directory--output-path (optional): Output path for merged datasets. Defaults to <dataset-path>/merged if not specifiedYour AIHub root directory should contain:
/path/to/aihub/
├── 027.일상생활 및 구어체 한-중, 한-일 번역 병렬 말뭉치 데이터/
├── 053.한국어-다국어(영어 제외) 번역 말뭉치(기술과학)/
├── 054.한국어-다국어 번역 말뭉치(기초과학)/
├── 055.한국어-다국어 번역 말뭉치(인문학)/
└── 한국어-일본어 번역 말뭉치/
The merged datasets will be saved to the specified output directory (default: output/) organized by language pair:
<output-path>/
├── ko_ja/ # Korean→Japanese translations
│ ├── train.json
│ ├── train.jsonl
│ ├── train.csv
│ ├── val.json
│ ├── val.jsonl
│ └── val.csv
├── ja_ko/ # Japanese→Korean translations
│ ├── train.json
│ ├── train.jsonl
│ ├── train.csv
│ ├── val.json
│ ├── val.jsonl
│ └── val.csv
└── ko_en/ # Korean→English translations
├── train.json
├── train.jsonl
├── train.csv
├── val.json
├── val.jsonl
└── val.csv
All formats contain the same data with sourceString and targetString fields:
sourceString,targetString)Each translation pair is saved as:
{
"sourceString": "source text in source language",
"targetString": "target text in target language"
}
The codebase is organized as follows:
AiHub/AiHubDatasetBase.py: Base class for all dataset implementationsAiHub/AiHubDataset*.py: Dataset-specific implementationsAiHub/util/DatasetGenerator.py: Utilities for merging and saving datasetsmain.py: Main script that orchestrates the merging processTo add support for a new AIHub dataset:
AiHub/ inheriting from AiHubDatasetBasemake_train_dataset(), make_val_dataset())source_lang and target_langmain.pySee CLAUDE.md for detailed architecture documentation.
MIT License
Python
100.0%
Code to merge some datasets in AIHub into single dataset
Python
2
15 commits
updated Nov 2, 2025
A Python tool for merging multiple AIHub translation datasets while preserving the original language direction. The tool processes various AIHub translation corpus datasets and outputs them grouped by language pair (source→target) in multiple formats.
The tool currently supports these AIHub datasets:
This project uses uv for dependency management:
# Install uv if you haven't already
pip install uv
# Clone the repository
git clone <repository-url>
cd aihub-translation-dataset
Run the main script with the path to your AIHub datasets:
uv run main.py --dataset-path /path/to/aihub
On Windows:
uv run main.py --dataset-path K:\DATASET\aihub
To specify a custom output directory:
uv run main.py --dataset-path /path/to/aihub --output-path /path/to/output
--dataset-path (required): Root path to the AIHub datasets directory--output-path (optional): Output path for merged datasets. Defaults to <dataset-path>/merged if not specifiedYour AIHub root directory should contain:
/path/to/aihub/
├── 027.일상생활 및 구어체 한-중, 한-일 번역 병렬 말뭉치 데이터/
├── 053.한국어-다국어(영어 제외) 번역 말뭉치(기술과학)/
├── 054.한국어-다국어 번역 말뭉치(기초과학)/
├── 055.한국어-다국어 번역 말뭉치(인문학)/
└── 한국어-일본어 번역 말뭉치/
The merged datasets will be saved to the specified output directory (default: output/) organized by language pair:
<output-path>/
├── ko_ja/ # Korean→Japanese translations
│ ├── train.json
│ ├── train.jsonl
│ ├── train.csv
│ ├── val.json
│ ├── val.jsonl
│ └── val.csv
├── ja_ko/ # Japanese→Korean translations
│ ├── train.json
│ ├── train.jsonl
│ ├── train.csv
│ ├── val.json
│ ├── val.jsonl
│ └── val.csv
└── ko_en/ # Korean→English translations
├── train.json
├── train.jsonl
├── train.csv
├── val.json
├── val.jsonl
└── val.csv
All formats contain the same data with sourceString and targetString fields:
sourceString,targetString)Each translation pair is saved as:
{
"sourceString": "source text in source language",
"targetString": "target text in target language"
}
The codebase is organized as follows:
AiHub/AiHubDatasetBase.py: Base class for all dataset implementationsAiHub/AiHubDataset*.py: Dataset-specific implementationsAiHub/util/DatasetGenerator.py: Utilities for merging and saving datasetsmain.py: Main script that orchestrates the merging processTo add support for a new AIHub dataset:
AiHub/ inheriting from AiHubDatasetBasemake_train_dataset(), make_val_dataset())source_lang and target_langmain.pySee CLAUDE.md for detailed architecture documentation.
MIT License
Python
100.0%