WildVision/wildvision-bench

Dataset

WildVision-Bench

11

17 commits

2 linked in READMEs

updated Oct 4, 2024

See the code

README

WildVision-Bench

We have two versions of Wildvision-Bench data

  • vision_bench_0617: the selected 500 examples that best simulates the vision-arena elo ranking, same data in the paper.
  • vision_bench_0701: the further filter and selected 500 examples by NSFW and manual selection. Leaderboard are still preparing.

Evaluation

Please refer to our Github for evaluation

If you want to evaluate your model, please use the vision_bench_0617 version to fairly compare the performance with other models in the following leaderboard.

Leaderboard (vision_bench_0717)

ModelScore95% CIWin RateRewardMuch BetterBetterTieWorseMuch WorseAvg Tokens
gpt-4o89.15(-1.9, 1.5)80.6%56.4255.0148.014.072.011.0142
gpt-4-vision-preview79.78(-2.9, 2.2)71.8%39.4182.0177.022.091.028.0138
Reka-Flash64.65(-2.6, 2.7)58.8%18.9135.0159.028.0116.062.0168
claude-3-opus-2024022962.03(-3.7, 2.8)53.0%13.5103.0162.048.0141.046.0105
yi-vl-plus55.05(-3.4, 2.3)52.8%7.298.0166.029.0124.083.0140
liuhaotian/llava-v1.6-34b51.89(-3.4, 3.8)49.2%2.590.0156.026.0145.083.0153
claude-3-sonnet-2024022950.0(0.0, 0.0)0.2%0.10.01.0499.00.00.0114
claude-3-haiku-2024030737.83(-2.6, 2.8)30.6%-16.554.099.047.0228.072.089
gemini-pro-vision35.57(-3.0, 3.2)32.6%-21.080.083.027.0167.0143.068
liuhaotian/llava-v1.6-vicuna-13b33.87(-2.9, 3.3)33.8%-21.462.0107.025.0167.0139.0136
deepseek-ai/deepseek-vl-7b-chat33.61(-3.3, 3.0)35.6%-21.259.0119.017.0161.0144.0116
THUDM/cogvlm-chat-hf32.01(-2.2, 3.0)30.6%-26.475.078.015.0172.0160.061
liuhaotian/llava-v1.6-vicuna-7b26.41(-3.3, 3.1)27.0%-31.445.090.036.0164.0165.0130
idefics2-8b-chatty23.96(-2.2, 2.4)26.4%-35.844.088.019.0164.0185.0135
Qwen/Qwen-VL-Chat18.08(-1.9, 2.2)19.6%-47.942.056.015.0155.0232.069
llava-1.5-7b-hf15.5(-2.4, 2.4)18.0%-47.828.062.025.0174.0211.0185
liuhaotian/llava-v1.5-13b14.43(-1.7, 1.6)16.8%-52.528.056.019.0157.0240.091
BAAI/Bunny-v1_0-3B12.98(-2.0, 2.1)16.6%-54.423.060.010.0164.0243.072
openbmb/MiniCPM-V11.95(-2.4, 2.1)13.6%-57.525.043.016.0164.0252.086
bczhou/tiny-llava-v1-hf8.3(-1.6, 1.2)11.0%-66.216.039.015.0127.0303.072
unum-cloud/uform-gen2-qwen-500m7.81(-1.3, 1.7)10.8%-68.516.038.011.0115.0320.092

Citation

@article{lu2024wildvision,
  title={WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences},
  author={Lu, Yujie and Jiang, Dongfu and Chen, Wenhu and Wang, William Yang and Choi, Yejin and Lin, Bill Yuchen},
  journal={arXiv preprint arXiv:2406.11069},
  year={2024}
}

Contributors

DongfuJiang

17 commits

WildVision/wildvision-bench

Dataset

WildVision-Bench

11

17 commits

2 linked in READMEs

updated Oct 4, 2024

See the code

README

WildVision-Bench

We have two versions of Wildvision-Bench data

  • vision_bench_0617: the selected 500 examples that best simulates the vision-arena elo ranking, same data in the paper.
  • vision_bench_0701: the further filter and selected 500 examples by NSFW and manual selection. Leaderboard are still preparing.

Evaluation

Please refer to our Github for evaluation

If you want to evaluate your model, please use the vision_bench_0617 version to fairly compare the performance with other models in the following leaderboard.

Leaderboard (vision_bench_0717)

ModelScore95% CIWin RateRewardMuch BetterBetterTieWorseMuch WorseAvg Tokens
gpt-4o89.15(-1.9, 1.5)80.6%56.4255.0148.014.072.011.0142
gpt-4-vision-preview79.78(-2.9, 2.2)71.8%39.4182.0177.022.091.028.0138
Reka-Flash64.65(-2.6, 2.7)58.8%18.9135.0159.028.0116.062.0168
claude-3-opus-2024022962.03(-3.7, 2.8)53.0%13.5103.0162.048.0141.046.0105
yi-vl-plus55.05(-3.4, 2.3)52.8%7.298.0166.029.0124.083.0140
liuhaotian/llava-v1.6-34b51.89(-3.4, 3.8)49.2%2.590.0156.026.0145.083.0153
claude-3-sonnet-2024022950.0(0.0, 0.0)0.2%0.10.01.0499.00.00.0114
claude-3-haiku-2024030737.83(-2.6, 2.8)30.6%-16.554.099.047.0228.072.089
gemini-pro-vision35.57(-3.0, 3.2)32.6%-21.080.083.027.0167.0143.068
liuhaotian/llava-v1.6-vicuna-13b33.87(-2.9, 3.3)33.8%-21.462.0107.025.0167.0139.0136
deepseek-ai/deepseek-vl-7b-chat33.61(-3.3, 3.0)35.6%-21.259.0119.017.0161.0144.0116
THUDM/cogvlm-chat-hf32.01(-2.2, 3.0)30.6%-26.475.078.015.0172.0160.061
liuhaotian/llava-v1.6-vicuna-7b26.41(-3.3, 3.1)27.0%-31.445.090.036.0164.0165.0130
idefics2-8b-chatty23.96(-2.2, 2.4)26.4%-35.844.088.019.0164.0185.0135
Qwen/Qwen-VL-Chat18.08(-1.9, 2.2)19.6%-47.942.056.015.0155.0232.069
llava-1.5-7b-hf15.5(-2.4, 2.4)18.0%-47.828.062.025.0174.0211.0185
liuhaotian/llava-v1.5-13b14.43(-1.7, 1.6)16.8%-52.528.056.019.0157.0240.091
BAAI/Bunny-v1_0-3B12.98(-2.0, 2.1)16.6%-54.423.060.010.0164.0243.072
openbmb/MiniCPM-V11.95(-2.4, 2.1)13.6%-57.525.043.016.0164.0252.086
bczhou/tiny-llava-v1-hf8.3(-1.6, 1.2)11.0%-66.216.039.015.0127.0303.072
unum-cloud/uform-gen2-qwen-500m7.81(-1.3, 1.7)10.8%-68.516.038.011.0115.0320.092

Citation

@article{lu2024wildvision,
  title={WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences},
  author={Lu, Yujie and Jiang, Dongfu and Chen, Wenhu and Wang, William Yang and Choi, Yejin and Lin, Bill Yuchen},
  journal={arXiv preprint arXiv:2406.11069},
  year={2024}
}

Contributors

DongfuJiang

17 commits