RekaAI/VibeEval

Dataset

Vibe-Eval

48

19 commits

2 linked in READMEs

updated Dec 12, 2024

See the code

README

Vibe-Eval

A benchmark for evaluating multimodal chat models, including especially challenging examples.

[Link to paper] [Blogpost] [Github]

Example from the dataset

Dataset

Each example has the following fields:

  • example_id: a unique ID for the example
  • category: the category that this example belongs to, either difficulty-normal or difficulty-hard
  • prompt: the user prompt
  • reference: a golden reference answer for the prompt
  • image: an image struct (containing bytes and path keys).
  • media_filename: the name of the file in the dataset
  • media_url: a URL where the file is hosted publicly

The dataset can also be downloaded from the Releases page of the reka-vibe-eval repo.

Leaderboard 🏆

Vibe-Eval Score (%)

Modelallhardnormal
Gemini Flash 2.067.152.375.9
Claude 3.5 Sonnet66.054.073.1
GPT-4o64.752.372.0
Gemini-1.5 Pro63.852.370.6
GPT-4o-mini56.744.763.8
Reka Flash56.039.3†65.8
Pixtral Large55.143.062.3
Grok Vision Beta54.237.164.2
Gemini 1.5 Flash 8b54.144.859.6
Claude Opus52.841.859.2
Pixtral 12b52.539.360.4
Claude Haiku48.531.658.2

† Note we expect the results of Reka models to be worse on the hard-set, as these are, by their very definition, prompts that Core cannot solve.

Running the evaluation

Check out github page to see instructions for evaluation.

Citation

@article{padlewski2024vibeeval,
  title={Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models},
  author={Piotr Padlewski and Max Bain and Matthew Henderson and Zhongkai Zhu and Nishant Relan and Hai Pham and Donovan Ong and Kaloyan Aleksiev and Aitor Ormazabal and Samuel Phua and Ethan Yeo and Eugenie Lamprecht and Qi Liu and Yuqi Wang and Eric Chen and Deyu Fu and Lei Li and Che Zheng and Cyprien de Masson d'Autume and Dani Yogatama and Mikel Artetxe and Yi Tay},
  journal={arXiv preprint arXiv:2405.02287},
  year={2024}
}
Eval
Hard
Reka
Vibe
Vibe-Eval
VibeEval

Contributors

prazek

14 commits

matthen

4 commits

m-bain

1 commits

RekaAI/VibeEval

Dataset

Vibe-Eval

48

19 commits

2 linked in READMEs

updated Dec 12, 2024

See the code

README

Vibe-Eval

A benchmark for evaluating multimodal chat models, including especially challenging examples.

[Link to paper] [Blogpost] [Github]

Example from the dataset

Dataset

Each example has the following fields:

  • example_id: a unique ID for the example
  • category: the category that this example belongs to, either difficulty-normal or difficulty-hard
  • prompt: the user prompt
  • reference: a golden reference answer for the prompt
  • image: an image struct (containing bytes and path keys).
  • media_filename: the name of the file in the dataset
  • media_url: a URL where the file is hosted publicly

The dataset can also be downloaded from the Releases page of the reka-vibe-eval repo.

Leaderboard 🏆

Vibe-Eval Score (%)

Modelallhardnormal
Gemini Flash 2.067.152.375.9
Claude 3.5 Sonnet66.054.073.1
GPT-4o64.752.372.0
Gemini-1.5 Pro63.852.370.6
GPT-4o-mini56.744.763.8
Reka Flash56.039.3†65.8
Pixtral Large55.143.062.3
Grok Vision Beta54.237.164.2
Gemini 1.5 Flash 8b54.144.859.6
Claude Opus52.841.859.2
Pixtral 12b52.539.360.4
Claude Haiku48.531.658.2

† Note we expect the results of Reka models to be worse on the hard-set, as these are, by their very definition, prompts that Core cannot solve.

Running the evaluation

Check out github page to see instructions for evaluation.

Citation

@article{padlewski2024vibeeval,
  title={Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models},
  author={Piotr Padlewski and Max Bain and Matthew Henderson and Zhongkai Zhu and Nishant Relan and Hai Pham and Donovan Ong and Kaloyan Aleksiev and Aitor Ormazabal and Samuel Phua and Ethan Yeo and Eugenie Lamprecht and Qi Liu and Yuqi Wang and Eric Chen and Deyu Fu and Lei Li and Che Zheng and Cyprien de Masson d'Autume and Dani Yogatama and Mikel Artetxe and Yi Tay},
  journal={arXiv preprint arXiv:2405.02287},
  year={2024}
}
Eval
Hard
Reka
Vibe
Vibe-Eval
VibeEval

Contributors

prazek

14 commits

matthen

4 commits

m-bain

1 commits