WeiChow/CrispEdit-2M

Dataset

17

stars

500

commits

1

linked in READMEs

Dec 29, 2025

updated

image
image-editing
instruction-guided
instruction-tuning
multimodal
Browse cluster: Image Editing and Manipulation

README

🖼️ CrispEdit-2M

CrispEdit-2M is a comprehensive dataset introduced in the paper ✨ EditMGT: Unleashing the Potential of Masked Generative Transformer in Image Editing ✨. This dataset encompasses 7 distinct image editing task categories.

arXiv Dataset Checkpoint GitHub Page

🌟 Overview

CrispEdit-2M is a large-scale dataset specifically designed for training and evaluating image editing models. With over 2.2 million samples across 7 different editing tasks, it provides researchers with a rich resource for developing advanced image manipulation techniques.

📊 Dataset Format

CrispEdit-2M contains 7 types of image editing tasks, stored in parquet files:

🏷️ Filename Prefix & Type in Parquet📝 Type Name🔢 Parquet Files (256 items per file)📈 Total Samples
colorColor Alteration1,984496K
motionMotion Change12832K
styleStyle Change1,600400K
replaceObject Replacement1,566391K
removeObject Removal1,388347K
addObject Addition1,213303K
backgroundBackground Change1,091272K
Total2,241K

Each parquet file in the CrispEdit-2M dataset contains 256 items, making it efficiently structured for large-scale image editing research.

🔍 Dataset Access

The complete dataset can be accessed through the Hugging Face repository. The dataset is organized by task categories for easy navigation and use.

from datasets import load_dataset

# Load the entire dataset
dataset = load_dataset("WeiChow/CrispEdit-2M")

📑 Citation

@article{chow2025editmgt,
  title={EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing},
  author={Chow, Wei and Li, Linfeng and Kong, Lingdong and Li, Zefeng and Xu, Qi and Song, Hang and Ye, Tian and Wang, Xian and Bai, Jinbin and Xu, Shilin and others},
  journal={arXiv preprint arXiv:2512.11715},
  year={2025}
}

🙏 Acknowledgements

We extend our sincere gratitude to all contributors and the research community for their valuable feedback and support in the development of this dataset.

Contributors

WeiChow

500 commits

WeiChow/CrispEdit-2M

Dataset

17

stars

500

commits

1

linked in READMEs

Dec 29, 2025

updated

image
image-editing
instruction-guided
instruction-tuning
multimodal
Browse cluster: Image Editing and Manipulation

README

🖼️ CrispEdit-2M

CrispEdit-2M is a comprehensive dataset introduced in the paper ✨ EditMGT: Unleashing the Potential of Masked Generative Transformer in Image Editing ✨. This dataset encompasses 7 distinct image editing task categories.

arXiv Dataset Checkpoint GitHub Page

🌟 Overview

CrispEdit-2M is a large-scale dataset specifically designed for training and evaluating image editing models. With over 2.2 million samples across 7 different editing tasks, it provides researchers with a rich resource for developing advanced image manipulation techniques.

📊 Dataset Format

CrispEdit-2M contains 7 types of image editing tasks, stored in parquet files:

🏷️ Filename Prefix & Type in Parquet📝 Type Name🔢 Parquet Files (256 items per file)📈 Total Samples
colorColor Alteration1,984496K
motionMotion Change12832K
styleStyle Change1,600400K
replaceObject Replacement1,566391K
removeObject Removal1,388347K
addObject Addition1,213303K
backgroundBackground Change1,091272K
Total2,241K

Each parquet file in the CrispEdit-2M dataset contains 256 items, making it efficiently structured for large-scale image editing research.

🔍 Dataset Access

The complete dataset can be accessed through the Hugging Face repository. The dataset is organized by task categories for easy navigation and use.

from datasets import load_dataset

# Load the entire dataset
dataset = load_dataset("WeiChow/CrispEdit-2M")

📑 Citation

@article{chow2025editmgt,
  title={EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing},
  author={Chow, Wei and Li, Linfeng and Kong, Lingdong and Li, Zefeng and Xu, Qi and Song, Hang and Ye, Tian and Wang, Xian and Bai, Jinbin and Xu, Shilin and others},
  journal={arXiv preprint arXiv:2512.11715},
  year={2025}
}

🙏 Acknowledgements

We extend our sincere gratitude to all contributors and the research community for their valuable feedback and support in the development of this dataset.

Contributors

WeiChow

500 commits