Rapidata/text-to-image-human-preference-6M

Dataset

Benchmark.ai Image Benchmark

10

6 commits

updated Oct 1, 2026

See the code

README

Benchmark.ai Image Benchmark

Built by Rapidata for Benchmark.ai.

This dataset contains 6,123,853 human responses, collected with the Rapidata Python SDK, comparing 43 text-to-image models on 1,514 prompts. Each row is a head-to-head comparison between two models' images for the same prompt, judged by human annotators on up to four questions β€” Alignment, Preference, Coherence and Text coherence.

Every image is ranked purely by human judgement, no automated metrics. This is the data behind the Benchmark.ai image leaderboard.

If you get value from this dataset and would like to see more in the future, please consider liking it ❀️

To evaluate your own models and create a leaderboard, check out our MRI.



πŸ† Live Leaderboard

Explore the full interactive leaderboard on Benchmark.ai β€” compare models side by side, browse matchups and prompts, and slice the standings by capability, output type, subject and prompt source.

Benchmark.ai Image Benchmark β€” click to open the interactive leaderboard

πŸ‘† Click to open it on benchmark.ai, or on Rapidata

At a glance

Head-to-head comparisons (rows)680,548
Human votes6,123,853 (Alignment 1,730,190 Β· Preference 1,731,352 Β· Coherence 1,730,112 Β· Text coherence 932,199)
Models compared43
Prompts1,514 (616 of them ask for rendered text)
Unique generated images64,266
Annotator countries148

The leaderboards

Every model is compared on the same kind of image pairs across four independent questions, run as four leaderboards on one Rapidata benchmark:

LeaderboardQuestion shown to annotatorsPrompt shown?Measures
Alignment"Which image matches the description better?"Yeshow faithfully the image depicts the prompt
Preference"Which image looks higher quality and more visually appealing?"Nooverall visual appeal, independent of the prompt
Coherence"Which image has more glitches and is more likely to be AI generated?"Novisual soundness β€” the image with fewer artifacts wins
Text coherence"In which image does the text have glitches or is harder to read?"Nolegibility of rendered text β€” the image with fewer text glitches wins (text prompts only)
.vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; } .image-container img { max-height: 300px; margin: 0; object-fit: contain; width: auto; box-sizing: content-box; } .image-container img.big { max-height: 400px; } .image-container { display: flex; justify-content: space-around; align-items: center; gap: .5rem } .container { width: 90%; margin: 0 auto; } .text-center { text-align: center; } .score-amount { margin: 0; margin-top: 10px; } .score-percentage { font-size: 12px; font-weight: semi-bold; }

Alignment

The alignment score measures how well the image matches its prompt. Annotators saw the prompt and were asked: "Which image matches the description better?"

a cat chasing a horse

Reve 2.1

Alignment ELO 1149 (#8) Β· 100% of votes

HunyuanImage 3.0

Alignment ELO 989 (#27) Β· 0% of votes
Book cover, deep navy background: a heraldic oval crest with a silver border, pierced by a katana and a crossbow bolt. Title (not touching the edge): "Comandante Vega: Orion".

GPT Image 2.5 (Flare)

Alignment ELO 1341 (#1) Β· 90% of votes

Pruna P-Image

Alignment ELO 634 (#43) Β· 10% of votes

Preference

The preference score reflects how visually appealing annotators found each image. Without seeing the prompt, they were asked: "Which image looks higher quality and more visually appealing?"

GPT Image 2.5 (Flare)

Preference ELO 1196 (#5) Β· 100% of votes

FLUX.2 klein 9B

Preference ELO 905 (#32) Β· 0% of votes

nano-banana-pro

Preference ELO 1044 (#19) Β· 100% of votes

Ideogram 4.0 Quality

Preference ELO 800 (#38) Β· 0% of votes

Coherence

The coherence score measures whether the image is free of artifacts and visual glitches. Without seeing the prompt, annotators were asked: "Which image has more glitches and is more likely to be AI generated?" The image picked less often wins.

Pruna P-Image

Coherence ELO 1130 (#6) Β· picked as glitchier by 0% of votes

Reve 2.1

Coherence ELO 1157 (#3) Β· picked as glitchier by 100% of votes

FLUX.2 max

Coherence ELO 975 (#25) Β· picked as glitchier by 0% of votes

Stable Diffusion 3.5 Large

Coherence ELO 862 (#41) Β· picked as glitchier by 100% of votes

Text coherence

The text coherence score measures how cleanly a model renders text. It runs only on the 616 prompts that ask for text. Without seeing the prompt, annotators were asked: "In which image does the text have glitches or is harder to read?" The image picked less often wins.

Seedream 4.5

Text coherence ELO 1220 (#1) Β· picked as glitchier by 0% of votes

Pruna P-Image

Text coherence ELO 884 (#33) Β· picked as glitchier by 100% of votes

Recraft V4.1 Utility Pro

Text coherence ELO 925 (#26) Β· picked as glitchier by 0% of votes

Pruna P-Image

Text coherence ELO 884 (#33) Β· picked as glitchier by 100% of votes

Rankings

Overall ranking (ELO, aggregated across all four leaderboards)

#ModelLabOverallAlignmentPreferenceCoherenceText coherence
1GPT Image 2.5 (Sunburst)OpenAI1238.31326125012461181
2GPT Image 2.5 (Flare)OpenAI1207.11341119612191136
3GPT Image 2.0OpenAI1167.31242113511131186
4Reve 2.1Reve AI1147.51149109611571193
5Grok Image Imagine 2.0 (Medium)SpaceXAI1136.81192109910961165
6GPT Image 1.5 (high)OpenAI1131.9124712428961154
7nano-banana-2Google1128.11169124210501051
8Grok Imagine Image 2.0SpaceXAI1107.51128108610761155
9Muse ImageMeta1096.31171110810621034
10Qwen-Image-3.0Alibaba Cloud1083.1111012059721031
11nano-banana-proGoogle1071.61070104411041049
12Seedream 4.5ByteDance1046.710329999481220
13Grok Image Imagine 2.0 (Thinking)SpaceXAI1045.61039108710121038
14MAI-Image 2.5Microsoft AI1042.1106011149751015
15Imagen 4 UltraGoogle1025.9105711068691081
16Seedream 5.0 LiteByteDance1013.710278969571183
17FLUX.2 flexBlack Forest Labs1011.0103410368721104
18Grok Imagine Image (quality)SpaceXAI1009.711391034880980
19Wan 2.7 ProAlibaba Cloud1007.610181038987981
20FLUX.2 proBlack Forest Labs1003.410159469901061
21Seedream 5.0 ProByteDance992.9100510761016858
22FLUX.2 maxBlack Forest Labs992.89929329751071
23Qwen-ImageAlibaba Cloud991.793311291003888
24Grok Imagine ImageSpaceXAI988.6103810168661031
25HiDream O1HiDream AI988.210431101871928
26Qwen-Image 2.0 ProAlibaba Cloud986.310059209391087
27Ideogram 4.5 QualityIdeogram967.685182111361006
28Recraft V4.1 Utility ProRecraft953.39838901004925
29FLUX.2 devBlack Forest Labs945.71016936911910
30Imagen 3Google939.39121014964823
31Imagen 4Google934.78791036901903
32Imagen 4 FastGoogle930.2893946953914
33FLUX.2 klein 9BBlack Forest Labs925.3950905992837
34HunyuanImage 3.0Tencent924.89891130767788
35Ideogram 4.0 QualityIdeogram911.98178001100911
36ERNIE ImageBaidu896.98601145694846
37Pruna-P-Ideogram (High)Pruna AI894.97677891101891
38Pruna P-ImagePruna AI890.26348741130884
39Pruna-P-Ideogram (Low)Pruna AI870.77697141105β€”
40Pruna-P-Ideogram (Medium)Pruna AI867.97446951137β€”
41Luma UNI 1.1 MaxLuma Labs846.18206191014872
42Pruna-P-Ideogram (Very Low)Pruna AI844.37536831076β€”
43Stable Diffusion 3.5 LargeStability AI794.4780868862627

A model shows "β€”" where it has no standing on that leaderboard.

By output type and capability

OpenAI's GPT Image 2.5 tops every category the benchmark slices by, so the last column shows the strongest model from any other lab. The full per-category standings, including subject and prompt source, are on Benchmark.ai.

Output typePrompts#1 overall (ELO)Best non-OpenAI model (ELO)
General Image524GPT Image 2.5 (Sunburst) (1265)Reve 2.1 β€” Reve AI (1151)
Photorealism243GPT Image 2.5 (Flare) (1242)Reve 2.1 β€” Reve AI (1141)
UI / UX182GPT Image 2.5 (Sunburst) (1213)Muse Image β€” Meta (1174)
Illustration & Anime147GPT Image 2.5 (Sunburst) (1237)Reve 2.1 β€” Reve AI (1161)
Graphic Design & Marketing137GPT Image 2.5 (Sunburst) (1249)Reve 2.1 β€” Reve AI (1222)
Logos & Branding134GPT Image 2.5 (Flare) (1209)nano-banana-2 β€” Google (1151)
Infographics131GPT Image 2.5 (Sunburst) (1235)nano-banana-2 β€” Google (1154)
3D Modeling & Rendering19GPT Image 2.5 (Sunburst) (1270)Qwen-Image-3.0 β€” Alibaba Cloud (1176)
CapabilityPrompts#1 overall (ELO)Best non-OpenAI model (ELO)
Attribute Binding758GPT Image 2.5 (Sunburst) (1214)Reve 2.1 β€” Reve AI (1153)
Text616GPT Image 2.5 (Sunburst) (1236)Reve 2.1 β€” Reve AI (1159)
Scene Composition428GPT Image 2.5 (Flare) (1230)Reve 2.1 β€” Reve AI (1163)
Knowledge & Reasoning384GPT Image 2.5 (Sunburst) (1227)nano-banana-2 β€” Google (1141)
Spatial & Size264GPT Image 2.5 (Sunburst) (1228)Reve 2.1 β€” Reve AI (1129)
Number of Objects252GPT Image 2.5 (Sunburst) (1265)Reve 2.1 β€” Reve AI (1169)
Surreal244GPT Image 2.5 (Flare) (1206)nano-banana-2 β€” Google (1142)
Emotion & Mood130GPT Image 2.5 (Sunburst) (1254)Reve 2.1 β€” Reve AI (1177)
Noisy Input20GPT Image 2.5 (Sunburst) (1377)Grok Image Imagine 2.0 (Medium) β€” SpaceXAI (1238)

How the dataset was built

1. Prompts (1,514)

The prompts are drawn from established text-to-image evaluation sets, real user prompts and prompts written for this benchmark:

SourcePrompts
Rapidata Original328
DiffusionDB325
PartiPrompts243
Community235
MONET187
DrawBench87
ABC-6K40
T2I-CompBench31
DALLE3-EVAL20
HRS-Bench18

Every prompt is tagged on three more axes, which the live leaderboard can slice by:

  • Capability β€” attribute binding, text, scene composition, knowledge & reasoning, spatial & size, number of objects, surreal, emotion & mood, noisy input (a prompt can carry several)
  • Output type β€” general image, photorealism, UI / UX, illustration & anime, graphic design & marketing, logos & branding, infographics, 3D modeling & rendering
  • Subject β€” objects & abstract, people & portraits, animals, architecture & spaces, vehicles, sci-fi & fantasy, food & drink, nature & landscapes

2. Image generation (43 models)

Each prompt was sent to every model, and the resulting images were uploaded to a Rapidata MRI benchmark as one participant per model. Every model has images for between 1,410 and 1,514 of the 1,514 prompts.

Alibaba Cloud β€” Qwen-Image, Qwen-Image 2.0 Pro, Qwen-Image-3.0, Wan 2.7 Pro Β· Baidu β€” ERNIE Image Β· Black Forest Labs β€” FLUX.2 dev, FLUX.2 flex, FLUX.2 klein 9B, FLUX.2 max, FLUX.2 pro Β· ByteDance β€” Seedream 4.5, Seedream 5.0 Lite, Seedream 5.0 Pro Β· Google β€” Imagen 3, Imagen 4, Imagen 4 Fast, Imagen 4 Ultra, nano-banana-2, nano-banana-pro Β· HiDream AI β€” HiDream O1 Β· Ideogram β€” Ideogram 4.0 Quality, Ideogram 4.5 Quality Β· Luma Labs β€” Luma UNI 1.1 Max Β· Meta β€” Muse Image Β· Microsoft AI β€” MAI-Image 2.5 Β· OpenAI β€” GPT Image 1.5 (high), GPT Image 2.0, GPT Image 2.5 (Flare), GPT Image 2.5 (Sunburst) Β· Pruna AI β€” Pruna P-Image, Pruna-P-Ideogram (High), Pruna-P-Ideogram (Low), Pruna-P-Ideogram (Medium), Pruna-P-Ideogram (Very Low) Β· Recraft β€” Recraft V4.1 Utility Pro Β· Reve AI β€” Reve 2.1 Β· SpaceXAI β€” Grok Image Imagine 2.0 (Medium), Grok Image Imagine 2.0 (Thinking), Grok Imagine Image, Grok Imagine Image (quality), Grok Imagine Image 2.0 Β· Stability AI β€” Stable Diffusion 3.5 Large Β· Tencent β€” HunyuanImage 3.0

3. Human evaluation

The four leaderboards were run over pairwise matchups between the models' images for the same prompt. Alignment, Coherence and Text coherence were answered by dedicated, qualified Rapidata audiences; Preference by the open global pool. Each pair is judged on a subset of the questions, with at least 5 votes per question it was judged on; Text coherence runs only on the 616 text prompts. Per-annotator detail β€” chosen side, country, language, gender, age bucket, occupation and the annotator's userScore β€” is preserved in the detailed_results_* columns. Votes came from annotators in 148 countries.

Dataset structure

One row per head-to-head image pair generated from the same prompt.

ColumnTypeDescription
promptstringthe prompt both images were generated from
image1stringpublic URL of the image from model1
image2stringpublic URL of the image from model2
model1stringname of the model that produced image1
model2stringname of the model that produced image2
weighted_results_image1_<q>float32sum of the userScore weights of the votes for image1 on question <q>
weighted_results_image2_<q>float32sum of the userScore weights of the votes for image2 on question <q>
detailed_results_<q>stringevery vote on question <q> as a JSON array (votedFor is A for image1 and B for image2, plus annotator demographics, userScore and votedAt)

<q> is one of alignment, preference, coherence and text_coherence. A row's columns for a question it was not judged on are empty.

  • The weighted_results_* values are userScore-weighted vote sums, not probabilities β€” they do not sum to 1. Divide by their sum for a normalised win share.
  • Coherence and Text coherence are already flipped. Annotators were asked which image has more glitches, but in this dataset votedFor and weighted_results_* point to the image with fewer glitches, so a higher value is better on every question.
import json
from datasets import load_dataset

ds = load_dataset("Rapidata/benchmark-ai-image-benchmark", split="train")
row = ds[0]
print(row["prompt"], "|", row["model1"], "vs", row["model2"])

for q in ["alignment", "preference", "coherence", "text_coherence"]:
    w1, w2 = row[f"weighted_results_image1_{q}"], row[f"weighted_results_image2_{q}"]
    if w1 is None or w1 + w2 == 0:
        continue  # this pair was not judged on q
    print(f"{q}: {row['model1']} wins {w1 / (w1 + w2):.0%} of the weighted vote")

Licensing & attribution

This dataset combines material under different terms:

  • Prompts. Prompts taken from public evaluation sets (PartiPrompts, DrawBench, DiffusionDB, MONET, T2I-CompBench, HRS-Bench, ABC-6K, DALLE3-EVAL) remain under their original licenses; Rapidata-original and community prompts are released under CC-BY-4.0.
  • Generated images. Produced by third-party models. Model outputs are governed by the terms of use of each respective model provider.
  • Human annotations. Collected via Rapidata, released under CC-BY-4.0.

About Rapidata

Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.

Explore all of our live benchmarks on Benchmark.ai.

alignment
benchmark
benchmark.ai
coherence
Human
image-generation
Preference
rapidata
text-to-image

Rapidata/text-to-image-human-preference-6M

Dataset

Benchmark.ai Image Benchmark

10

6 commits

updated Oct 1, 2026

See the code

README

Benchmark.ai Image Benchmark

Built by Rapidata for Benchmark.ai.

This dataset contains 6,123,853 human responses, collected with the Rapidata Python SDK, comparing 43 text-to-image models on 1,514 prompts. Each row is a head-to-head comparison between two models' images for the same prompt, judged by human annotators on up to four questions β€” Alignment, Preference, Coherence and Text coherence.

Every image is ranked purely by human judgement, no automated metrics. This is the data behind the Benchmark.ai image leaderboard.

If you get value from this dataset and would like to see more in the future, please consider liking it ❀️

To evaluate your own models and create a leaderboard, check out our MRI.



πŸ† Live Leaderboard

Explore the full interactive leaderboard on Benchmark.ai β€” compare models side by side, browse matchups and prompts, and slice the standings by capability, output type, subject and prompt source.

Benchmark.ai Image Benchmark β€” click to open the interactive leaderboard

πŸ‘† Click to open it on benchmark.ai, or on Rapidata

At a glance

Head-to-head comparisons (rows)680,548
Human votes6,123,853 (Alignment 1,730,190 Β· Preference 1,731,352 Β· Coherence 1,730,112 Β· Text coherence 932,199)
Models compared43
Prompts1,514 (616 of them ask for rendered text)
Unique generated images64,266
Annotator countries148

The leaderboards

Every model is compared on the same kind of image pairs across four independent questions, run as four leaderboards on one Rapidata benchmark:

LeaderboardQuestion shown to annotatorsPrompt shown?Measures
Alignment"Which image matches the description better?"Yeshow faithfully the image depicts the prompt
Preference"Which image looks higher quality and more visually appealing?"Nooverall visual appeal, independent of the prompt
Coherence"Which image has more glitches and is more likely to be AI generated?"Novisual soundness β€” the image with fewer artifacts wins
Text coherence"In which image does the text have glitches or is harder to read?"Nolegibility of rendered text β€” the image with fewer text glitches wins (text prompts only)
.vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; } .image-container img { max-height: 300px; margin: 0; object-fit: contain; width: auto; box-sizing: content-box; } .image-container img.big { max-height: 400px; } .image-container { display: flex; justify-content: space-around; align-items: center; gap: .5rem } .container { width: 90%; margin: 0 auto; } .text-center { text-align: center; } .score-amount { margin: 0; margin-top: 10px; } .score-percentage { font-size: 12px; font-weight: semi-bold; }

Alignment

The alignment score measures how well the image matches its prompt. Annotators saw the prompt and were asked: "Which image matches the description better?"

a cat chasing a horse

Reve 2.1

Alignment ELO 1149 (#8) Β· 100% of votes

HunyuanImage 3.0

Alignment ELO 989 (#27) Β· 0% of votes
Book cover, deep navy background: a heraldic oval crest with a silver border, pierced by a katana and a crossbow bolt. Title (not touching the edge): "Comandante Vega: Orion".

GPT Image 2.5 (Flare)

Alignment ELO 1341 (#1) Β· 90% of votes

Pruna P-Image

Alignment ELO 634 (#43) Β· 10% of votes

Preference

The preference score reflects how visually appealing annotators found each image. Without seeing the prompt, they were asked: "Which image looks higher quality and more visually appealing?"

GPT Image 2.5 (Flare)

Preference ELO 1196 (#5) Β· 100% of votes

FLUX.2 klein 9B

Preference ELO 905 (#32) Β· 0% of votes

nano-banana-pro

Preference ELO 1044 (#19) Β· 100% of votes

Ideogram 4.0 Quality

Preference ELO 800 (#38) Β· 0% of votes

Coherence

The coherence score measures whether the image is free of artifacts and visual glitches. Without seeing the prompt, annotators were asked: "Which image has more glitches and is more likely to be AI generated?" The image picked less often wins.

Pruna P-Image

Coherence ELO 1130 (#6) Β· picked as glitchier by 0% of votes

Reve 2.1

Coherence ELO 1157 (#3) Β· picked as glitchier by 100% of votes

FLUX.2 max

Coherence ELO 975 (#25) Β· picked as glitchier by 0% of votes

Stable Diffusion 3.5 Large

Coherence ELO 862 (#41) Β· picked as glitchier by 100% of votes

Text coherence

The text coherence score measures how cleanly a model renders text. It runs only on the 616 prompts that ask for text. Without seeing the prompt, annotators were asked: "In which image does the text have glitches or is harder to read?" The image picked less often wins.

Seedream 4.5

Text coherence ELO 1220 (#1) Β· picked as glitchier by 0% of votes

Pruna P-Image

Text coherence ELO 884 (#33) Β· picked as glitchier by 100% of votes

Recraft V4.1 Utility Pro

Text coherence ELO 925 (#26) Β· picked as glitchier by 0% of votes

Pruna P-Image

Text coherence ELO 884 (#33) Β· picked as glitchier by 100% of votes

Rankings

Overall ranking (ELO, aggregated across all four leaderboards)

#ModelLabOverallAlignmentPreferenceCoherenceText coherence
1GPT Image 2.5 (Sunburst)OpenAI1238.31326125012461181
2GPT Image 2.5 (Flare)OpenAI1207.11341119612191136
3GPT Image 2.0OpenAI1167.31242113511131186
4Reve 2.1Reve AI1147.51149109611571193
5Grok Image Imagine 2.0 (Medium)SpaceXAI1136.81192109910961165
6GPT Image 1.5 (high)OpenAI1131.9124712428961154
7nano-banana-2Google1128.11169124210501051
8Grok Imagine Image 2.0SpaceXAI1107.51128108610761155
9Muse ImageMeta1096.31171110810621034
10Qwen-Image-3.0Alibaba Cloud1083.1111012059721031
11nano-banana-proGoogle1071.61070104411041049
12Seedream 4.5ByteDance1046.710329999481220
13Grok Image Imagine 2.0 (Thinking)SpaceXAI1045.61039108710121038
14MAI-Image 2.5Microsoft AI1042.1106011149751015
15Imagen 4 UltraGoogle1025.9105711068691081
16Seedream 5.0 LiteByteDance1013.710278969571183
17FLUX.2 flexBlack Forest Labs1011.0103410368721104
18Grok Imagine Image (quality)SpaceXAI1009.711391034880980
19Wan 2.7 ProAlibaba Cloud1007.610181038987981
20FLUX.2 proBlack Forest Labs1003.410159469901061
21Seedream 5.0 ProByteDance992.9100510761016858
22FLUX.2 maxBlack Forest Labs992.89929329751071
23Qwen-ImageAlibaba Cloud991.793311291003888
24Grok Imagine ImageSpaceXAI988.6103810168661031
25HiDream O1HiDream AI988.210431101871928
26Qwen-Image 2.0 ProAlibaba Cloud986.310059209391087
27Ideogram 4.5 QualityIdeogram967.685182111361006
28Recraft V4.1 Utility ProRecraft953.39838901004925
29FLUX.2 devBlack Forest Labs945.71016936911910
30Imagen 3Google939.39121014964823
31Imagen 4Google934.78791036901903
32Imagen 4 FastGoogle930.2893946953914
33FLUX.2 klein 9BBlack Forest Labs925.3950905992837
34HunyuanImage 3.0Tencent924.89891130767788
35Ideogram 4.0 QualityIdeogram911.98178001100911
36ERNIE ImageBaidu896.98601145694846
37Pruna-P-Ideogram (High)Pruna AI894.97677891101891
38Pruna P-ImagePruna AI890.26348741130884
39Pruna-P-Ideogram (Low)Pruna AI870.77697141105β€”
40Pruna-P-Ideogram (Medium)Pruna AI867.97446951137β€”
41Luma UNI 1.1 MaxLuma Labs846.18206191014872
42Pruna-P-Ideogram (Very Low)Pruna AI844.37536831076β€”
43Stable Diffusion 3.5 LargeStability AI794.4780868862627

A model shows "β€”" where it has no standing on that leaderboard.

By output type and capability

OpenAI's GPT Image 2.5 tops every category the benchmark slices by, so the last column shows the strongest model from any other lab. The full per-category standings, including subject and prompt source, are on Benchmark.ai.

Output typePrompts#1 overall (ELO)Best non-OpenAI model (ELO)
General Image524GPT Image 2.5 (Sunburst) (1265)Reve 2.1 β€” Reve AI (1151)
Photorealism243GPT Image 2.5 (Flare) (1242)Reve 2.1 β€” Reve AI (1141)
UI / UX182GPT Image 2.5 (Sunburst) (1213)Muse Image β€” Meta (1174)
Illustration & Anime147GPT Image 2.5 (Sunburst) (1237)Reve 2.1 β€” Reve AI (1161)
Graphic Design & Marketing137GPT Image 2.5 (Sunburst) (1249)Reve 2.1 β€” Reve AI (1222)
Logos & Branding134GPT Image 2.5 (Flare) (1209)nano-banana-2 β€” Google (1151)
Infographics131GPT Image 2.5 (Sunburst) (1235)nano-banana-2 β€” Google (1154)
3D Modeling & Rendering19GPT Image 2.5 (Sunburst) (1270)Qwen-Image-3.0 β€” Alibaba Cloud (1176)
CapabilityPrompts#1 overall (ELO)Best non-OpenAI model (ELO)
Attribute Binding758GPT Image 2.5 (Sunburst) (1214)Reve 2.1 β€” Reve AI (1153)
Text616GPT Image 2.5 (Sunburst) (1236)Reve 2.1 β€” Reve AI (1159)
Scene Composition428GPT Image 2.5 (Flare) (1230)Reve 2.1 β€” Reve AI (1163)
Knowledge & Reasoning384GPT Image 2.5 (Sunburst) (1227)nano-banana-2 β€” Google (1141)
Spatial & Size264GPT Image 2.5 (Sunburst) (1228)Reve 2.1 β€” Reve AI (1129)
Number of Objects252GPT Image 2.5 (Sunburst) (1265)Reve 2.1 β€” Reve AI (1169)
Surreal244GPT Image 2.5 (Flare) (1206)nano-banana-2 β€” Google (1142)
Emotion & Mood130GPT Image 2.5 (Sunburst) (1254)Reve 2.1 β€” Reve AI (1177)
Noisy Input20GPT Image 2.5 (Sunburst) (1377)Grok Image Imagine 2.0 (Medium) β€” SpaceXAI (1238)

How the dataset was built

1. Prompts (1,514)

The prompts are drawn from established text-to-image evaluation sets, real user prompts and prompts written for this benchmark:

SourcePrompts
Rapidata Original328
DiffusionDB325
PartiPrompts243
Community235
MONET187
DrawBench87
ABC-6K40
T2I-CompBench31
DALLE3-EVAL20
HRS-Bench18

Every prompt is tagged on three more axes, which the live leaderboard can slice by:

  • Capability β€” attribute binding, text, scene composition, knowledge & reasoning, spatial & size, number of objects, surreal, emotion & mood, noisy input (a prompt can carry several)
  • Output type β€” general image, photorealism, UI / UX, illustration & anime, graphic design & marketing, logos & branding, infographics, 3D modeling & rendering
  • Subject β€” objects & abstract, people & portraits, animals, architecture & spaces, vehicles, sci-fi & fantasy, food & drink, nature & landscapes

2. Image generation (43 models)

Each prompt was sent to every model, and the resulting images were uploaded to a Rapidata MRI benchmark as one participant per model. Every model has images for between 1,410 and 1,514 of the 1,514 prompts.

Alibaba Cloud β€” Qwen-Image, Qwen-Image 2.0 Pro, Qwen-Image-3.0, Wan 2.7 Pro Β· Baidu β€” ERNIE Image Β· Black Forest Labs β€” FLUX.2 dev, FLUX.2 flex, FLUX.2 klein 9B, FLUX.2 max, FLUX.2 pro Β· ByteDance β€” Seedream 4.5, Seedream 5.0 Lite, Seedream 5.0 Pro Β· Google β€” Imagen 3, Imagen 4, Imagen 4 Fast, Imagen 4 Ultra, nano-banana-2, nano-banana-pro Β· HiDream AI β€” HiDream O1 Β· Ideogram β€” Ideogram 4.0 Quality, Ideogram 4.5 Quality Β· Luma Labs β€” Luma UNI 1.1 Max Β· Meta β€” Muse Image Β· Microsoft AI β€” MAI-Image 2.5 Β· OpenAI β€” GPT Image 1.5 (high), GPT Image 2.0, GPT Image 2.5 (Flare), GPT Image 2.5 (Sunburst) Β· Pruna AI β€” Pruna P-Image, Pruna-P-Ideogram (High), Pruna-P-Ideogram (Low), Pruna-P-Ideogram (Medium), Pruna-P-Ideogram (Very Low) Β· Recraft β€” Recraft V4.1 Utility Pro Β· Reve AI β€” Reve 2.1 Β· SpaceXAI β€” Grok Image Imagine 2.0 (Medium), Grok Image Imagine 2.0 (Thinking), Grok Imagine Image, Grok Imagine Image (quality), Grok Imagine Image 2.0 Β· Stability AI β€” Stable Diffusion 3.5 Large Β· Tencent β€” HunyuanImage 3.0

3. Human evaluation

The four leaderboards were run over pairwise matchups between the models' images for the same prompt. Alignment, Coherence and Text coherence were answered by dedicated, qualified Rapidata audiences; Preference by the open global pool. Each pair is judged on a subset of the questions, with at least 5 votes per question it was judged on; Text coherence runs only on the 616 text prompts. Per-annotator detail β€” chosen side, country, language, gender, age bucket, occupation and the annotator's userScore β€” is preserved in the detailed_results_* columns. Votes came from annotators in 148 countries.

Dataset structure

One row per head-to-head image pair generated from the same prompt.

ColumnTypeDescription
promptstringthe prompt both images were generated from
image1stringpublic URL of the image from model1
image2stringpublic URL of the image from model2
model1stringname of the model that produced image1
model2stringname of the model that produced image2
weighted_results_image1_<q>float32sum of the userScore weights of the votes for image1 on question <q>
weighted_results_image2_<q>float32sum of the userScore weights of the votes for image2 on question <q>
detailed_results_<q>stringevery vote on question <q> as a JSON array (votedFor is A for image1 and B for image2, plus annotator demographics, userScore and votedAt)

<q> is one of alignment, preference, coherence and text_coherence. A row's columns for a question it was not judged on are empty.

  • The weighted_results_* values are userScore-weighted vote sums, not probabilities β€” they do not sum to 1. Divide by their sum for a normalised win share.
  • Coherence and Text coherence are already flipped. Annotators were asked which image has more glitches, but in this dataset votedFor and weighted_results_* point to the image with fewer glitches, so a higher value is better on every question.
import json
from datasets import load_dataset

ds = load_dataset("Rapidata/benchmark-ai-image-benchmark", split="train")
row = ds[0]
print(row["prompt"], "|", row["model1"], "vs", row["model2"])

for q in ["alignment", "preference", "coherence", "text_coherence"]:
    w1, w2 = row[f"weighted_results_image1_{q}"], row[f"weighted_results_image2_{q}"]
    if w1 is None or w1 + w2 == 0:
        continue  # this pair was not judged on q
    print(f"{q}: {row['model1']} wins {w1 / (w1 + w2):.0%} of the weighted vote")

Licensing & attribution

This dataset combines material under different terms:

  • Prompts. Prompts taken from public evaluation sets (PartiPrompts, DrawBench, DiffusionDB, MONET, T2I-CompBench, HRS-Bench, ABC-6K, DALLE3-EVAL) remain under their original licenses; Rapidata-original and community prompts are released under CC-BY-4.0.
  • Generated images. Produced by third-party models. Model outputs are governed by the terms of use of each respective model provider.
  • Human annotations. Collected via Rapidata, released under CC-BY-4.0.

About Rapidata

Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.

Explore all of our live benchmarks on Benchmark.ai.

alignment
benchmark
benchmark.ai
coherence
Human
image-generation
Preference
rapidata
text-to-image