ConceptEdit-12M is a large-scale image editing dataset. Each sample is stored as a triplet:
The dataset is packaged as multiple .tar shards. All paths inside the tar files and JSON files are relative paths; no local absolute paths are included.
The released data contains four main splits:
| Split | Number of tar shards |
|---|---|
enhanced_prompt_square_resolution | 154 |
enhanced_prompt_random_resolution | 162 |
original_prompt_square_resolution | 131 |
original_prompt_random_resolution | 169 |
After downloading, the dataset is organized as follows:
ConceptEdit-12M/
├── enhanced_prompt_square_resolution/
│ ├── batch_0.tar
│ ├── batch_1.tar
│ └── ...
├── enhanced_prompt_random_resolution/
│ ├── batch_0.tar
│ ├── batch_1.tar
│ └── ...
├── original_prompt_square_resolution/
│ ├── batch_0.tar
│ ├── batch_1.tar
│ └── ...
└── original_prompt_random_resolution/
├── batch_0.tar
├── batch_1.tar
└── ...
Each tar shard contains one batch directory. For example:
batch_0/
├── 0_0_3.json
├── 0_0_3_source.jpg
├── 0_0_3_edit.png
├── 0_0_4.json
├── 0_0_4_source.jpg
├── 0_0_4_edit.png
└── ...
For each sample ID, the three files are:
<sample_id>.json
<sample_id>_source.jpg
<sample_id>_edit.png
The JSON file contains relative paths to the corresponding source and edited images, for example:
{
"id": "0_0_3",
"images": {
"source": "batch_0/0_0_3_source.jpg",
"edited": "batch_0/0_0_3_edit.png"
}
}
The following commands are provided for format reference only. You may use faster extraction tools or multi-threaded extraction methods according to your own needs.
Extract a single shard:
mkdir -p ConceptEdit-12M/extracted/enhanced_prompt_square_resolution
tar -xf ConceptEdit-12M/enhanced_prompt_square_resolution/batch_0.tar \
-C ConceptEdit-12M/extracted/enhanced_prompt_square_resolution
List the content of a shard without extracting:
tar -tf ConceptEdit-12M/enhanced_prompt_square_resolution/batch_0.tar | head
Extract all shards in one split:
SPLIT=enhanced_prompt_square_resolution
mkdir -p ConceptEdit-12M/extracted/${SPLIT}
for f in ConceptEdit-12M/${SPLIT}/batch_*.tar; do
tar -xf "$f" -C ConceptEdit-12M/extracted/${SPLIT}
done
Extract all four main splits:
for SPLIT in \
enhanced_prompt_square_resolution \
enhanced_prompt_random_resolution \
original_prompt_square_resolution \
original_prompt_random_resolution; do
mkdir -p ConceptEdit-12M/extracted/${SPLIT}
for f in ConceptEdit-12M/${SPLIT}/batch_*.tar; do
tar -xf "$f" -C ConceptEdit-12M/extracted/${SPLIT}
done
done
Each sample has one JSON file with the following top-level fields:
id
images
edit_concept
instruction
evaluation
original_simple_caption
Field descriptions:
id: sample ID. The image and JSON filenames are derived from this ID.images: relative paths to the source image and edited image.
source: source/original image path inside the extracted batch directory.edited: edited image path inside the extracted batch directory.edit_concept: taxonomy information for the edit.
categorysub_categorytaskdetailinstruction: natural-language edit instructions.
short_enshort_zhdetailed_endetailed_zhevaluation: VQA-style filtering and quality-control metadata.
overall_vqa_scorekeepwrong_countrecaption_prompt_enrecaption_prompt_zhvqa: a list of VQA checks, each containing:
dimensionquestion_enquestion_zhexpected_answerpassedoriginal_simple_caption: a short English caption for the source image.A simplified example:
{
"id": "0_0_3",
"images": {
"source": "batch_0/0_0_3_source.jpg",
"edited": "batch_0/0_0_3_edit.png"
},
"edit_concept": {
"category": "...",
"sub_category": "...",
"task": "...",
"detail": "..."
},
"instruction": {
"short_en": "...",
"short_zh": "...",
"detailed_en": "...",
"detailed_zh": "..."
},
"evaluation": {
"overall_vqa_score": 1.0,
"keep": true,
"wrong_count": 0,
"recaption_prompt_en": "",
"recaption_prompt_zh": "",
"vqa": [
{
"dimension": "1_primary_edit_success",
"question_en": "...",
"question_zh": "...",
"expected_answer": "yes",
"passed": true
}
]
},
"original_simple_caption": "A short English caption for the source image."
}
The source images in ConceptEdit-12M are based on images from Fine-T2I.
If you find this dataset useful, please cite:
@article{cui2026unlocking,
title={Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision},
author={Cui, Long and Liu, Xiaoqian and Qin, Qi and Xin, Yi and Lin, Tao and Li, Jianguo and Zhang, Linfeng},
journal={arXiv preprint arXiv:2608.16812},
year={2026}
}
500 commits
ConceptEdit-12M is a large-scale image editing dataset. Each sample is stored as a triplet:
The dataset is packaged as multiple .tar shards. All paths inside the tar files and JSON files are relative paths; no local absolute paths are included.
The released data contains four main splits:
| Split | Number of tar shards |
|---|---|
enhanced_prompt_square_resolution | 154 |
enhanced_prompt_random_resolution | 162 |
original_prompt_square_resolution | 131 |
original_prompt_random_resolution | 169 |
After downloading, the dataset is organized as follows:
ConceptEdit-12M/
├── enhanced_prompt_square_resolution/
│ ├── batch_0.tar
│ ├── batch_1.tar
│ └── ...
├── enhanced_prompt_random_resolution/
│ ├── batch_0.tar
│ ├── batch_1.tar
│ └── ...
├── original_prompt_square_resolution/
│ ├── batch_0.tar
│ ├── batch_1.tar
│ └── ...
└── original_prompt_random_resolution/
├── batch_0.tar
├── batch_1.tar
└── ...
Each tar shard contains one batch directory. For example:
batch_0/
├── 0_0_3.json
├── 0_0_3_source.jpg
├── 0_0_3_edit.png
├── 0_0_4.json
├── 0_0_4_source.jpg
├── 0_0_4_edit.png
└── ...
For each sample ID, the three files are:
<sample_id>.json
<sample_id>_source.jpg
<sample_id>_edit.png
The JSON file contains relative paths to the corresponding source and edited images, for example:
{
"id": "0_0_3",
"images": {
"source": "batch_0/0_0_3_source.jpg",
"edited": "batch_0/0_0_3_edit.png"
}
}
The following commands are provided for format reference only. You may use faster extraction tools or multi-threaded extraction methods according to your own needs.
Extract a single shard:
mkdir -p ConceptEdit-12M/extracted/enhanced_prompt_square_resolution
tar -xf ConceptEdit-12M/enhanced_prompt_square_resolution/batch_0.tar \
-C ConceptEdit-12M/extracted/enhanced_prompt_square_resolution
List the content of a shard without extracting:
tar -tf ConceptEdit-12M/enhanced_prompt_square_resolution/batch_0.tar | head
Extract all shards in one split:
SPLIT=enhanced_prompt_square_resolution
mkdir -p ConceptEdit-12M/extracted/${SPLIT}
for f in ConceptEdit-12M/${SPLIT}/batch_*.tar; do
tar -xf "$f" -C ConceptEdit-12M/extracted/${SPLIT}
done
Extract all four main splits:
for SPLIT in \
enhanced_prompt_square_resolution \
enhanced_prompt_random_resolution \
original_prompt_square_resolution \
original_prompt_random_resolution; do
mkdir -p ConceptEdit-12M/extracted/${SPLIT}
for f in ConceptEdit-12M/${SPLIT}/batch_*.tar; do
tar -xf "$f" -C ConceptEdit-12M/extracted/${SPLIT}
done
done
Each sample has one JSON file with the following top-level fields:
id
images
edit_concept
instruction
evaluation
original_simple_caption
Field descriptions:
id: sample ID. The image and JSON filenames are derived from this ID.images: relative paths to the source image and edited image.
source: source/original image path inside the extracted batch directory.edited: edited image path inside the extracted batch directory.edit_concept: taxonomy information for the edit.
categorysub_categorytaskdetailinstruction: natural-language edit instructions.
short_enshort_zhdetailed_endetailed_zhevaluation: VQA-style filtering and quality-control metadata.
overall_vqa_scorekeepwrong_countrecaption_prompt_enrecaption_prompt_zhvqa: a list of VQA checks, each containing:
dimensionquestion_enquestion_zhexpected_answerpassedoriginal_simple_caption: a short English caption for the source image.A simplified example:
{
"id": "0_0_3",
"images": {
"source": "batch_0/0_0_3_source.jpg",
"edited": "batch_0/0_0_3_edit.png"
},
"edit_concept": {
"category": "...",
"sub_category": "...",
"task": "...",
"detail": "..."
},
"instruction": {
"short_en": "...",
"short_zh": "...",
"detailed_en": "...",
"detailed_zh": "..."
},
"evaluation": {
"overall_vqa_score": 1.0,
"keep": true,
"wrong_count": 0,
"recaption_prompt_en": "",
"recaption_prompt_zh": "",
"vqa": [
{
"dimension": "1_primary_edit_success",
"question_en": "...",
"question_zh": "...",
"expected_answer": "yes",
"passed": true
}
]
},
"original_simple_caption": "A short English caption for the source image."
}
The source images in ConceptEdit-12M are based on images from Fine-T2I.
If you find this dataset useful, please cite:
@article{cui2026unlocking,
title={Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision},
author={Cui, Long and Liu, Xiaoqian and Qin, Qi and Xin, Yi and Lin, Tao and Li, Jianguo and Zhang, Linfeng},
journal={arXiv preprint arXiv:2608.16812},
year={2026}
}
500 commits