Dataset Card for GraphicDesignEvaluation
1
103 commits
1 linked in READMEs
updated Jun 26, 2026
GraphicDesignEvaluation is a human-rated benchmark released with Can GPTs Evaluate Graphic Design Based on Design Principles?. The paper compares GPT-based evaluation and heuristic metrics against human ratings for three representative design principles: alignment, overlap, and white space. The dataset contains graphic banner designs curated from an online service, perturbed low-quality variants, and human annotations collected from 60 subjects.
The dataset supports graphic design quality evaluation, human/model score correlation analysis, and design-principle-specific assessment. No public leaderboard is bundled with this Hugging Face dataset.
Annotations and evaluation descriptions are in English (en).
Absolute configs contain image_id, image, perturbation, scores, and avg. Relative configs contain image_id, image, comparative, scores, and avg.
All configs expose a single train split. Absolute configs have 400 rows each; relative configs have 300 rows each.
The dataset was created to study whether GPT-based evaluators can assess graphic design quality according to core design principles and how those scores compare with human annotations.
The dataset is small and principle-specific. It should be used as an evaluation resource rather than a complete measure of graphic design quality.
The local loader lists the dataset license as Apache 2.0.
@inproceedings{haraguchi2024can,
title={Can GPTs Evaluate Graphic Design Based on Design Principles?},
author={Haraguchi, Daichi and Inoue, Naoto and Shimoda, Wataru and Mitani, Hayato and Uchida, Seiichi and Yamaguchi, Kota},
booktitle={SIGGRAPH Asia 2024 Technical Communications},
pages={1--4},
year={2024}
}
103 commits
Dataset Card for GraphicDesignEvaluation
1
103 commits
1 linked in READMEs
updated Jun 26, 2026
GraphicDesignEvaluation is a human-rated benchmark released with Can GPTs Evaluate Graphic Design Based on Design Principles?. The paper compares GPT-based evaluation and heuristic metrics against human ratings for three representative design principles: alignment, overlap, and white space. The dataset contains graphic banner designs curated from an online service, perturbed low-quality variants, and human annotations collected from 60 subjects.
The dataset supports graphic design quality evaluation, human/model score correlation analysis, and design-principle-specific assessment. No public leaderboard is bundled with this Hugging Face dataset.
Annotations and evaluation descriptions are in English (en).
Absolute configs contain image_id, image, perturbation, scores, and avg. Relative configs contain image_id, image, comparative, scores, and avg.
All configs expose a single train split. Absolute configs have 400 rows each; relative configs have 300 rows each.
The dataset was created to study whether GPT-based evaluators can assess graphic design quality according to core design principles and how those scores compare with human annotations.
The dataset is small and principle-specific. It should be used as an evaluation resource rather than a complete measure of graphic design quality.
The local loader lists the dataset license as Apache 2.0.
@inproceedings{haraguchi2024can,
title={Can GPTs Evaluate Graphic Design Based on Design Principles?},
author={Haraguchi, Daichi and Inoue, Naoto and Shimoda, Wataru and Mitani, Hayato and Uchida, Seiichi and Yamaguchi, Kota},
booktitle={SIGGRAPH Asia 2024 Technical Communications},
pages={1--4},
year={2024}
}
103 commits