AILab-CVC/SEED-Bench-H

Dataset

3

stars

11

commits

10

linked in READMEs

May 30, 2024

updated

README


license: cc-by-nc-4.0 task_categories:

  • visual-question-answering language:
  • en pretty_name: SEED-Bench-H size_categories:
  • 1K<n<10K

SEED-Bench-H Card

Benchmark details

Benchmark type: SEED-Bench-H is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs). It consists of 28K multiple-choice questions with precise human annotations, spanning 34 dimensions, including the evaluation of both text and image generation.

Benchmark date: SEED-Bench-H was collected in April 2024.

Paper or resources for more information: https://github.com/AILab-CVC/SEED-Bench

License: Attribution-NonCommercial 4.0 International. It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use.

Data Sources:

Where to send questions or comments about the benchmark: https://github.com/AILab-CVC/SEED-Bench/issues

Intended use

Primary intended uses: The primary use of SEED-Bench-H is evaluate Multimodal Large Language Models in text and image generation tasks.

Primary intended users: The primary intended users of the Benchmark are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.

Contributors

BreakLee

11 commits

AILab-CVC/SEED-Bench-H

Dataset

3

stars

11

commits

10

linked in READMEs

May 30, 2024

updated

README


license: cc-by-nc-4.0 task_categories:

  • visual-question-answering language:
  • en pretty_name: SEED-Bench-H size_categories:
  • 1K<n<10K

SEED-Bench-H Card

Benchmark details

Benchmark type: SEED-Bench-H is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs). It consists of 28K multiple-choice questions with precise human annotations, spanning 34 dimensions, including the evaluation of both text and image generation.

Benchmark date: SEED-Bench-H was collected in April 2024.

Paper or resources for more information: https://github.com/AILab-CVC/SEED-Bench

License: Attribution-NonCommercial 4.0 International. It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use.

Data Sources:

Where to send questions or comments about the benchmark: https://github.com/AILab-CVC/SEED-Bench/issues

Intended use

Primary intended uses: The primary use of SEED-Bench-H is evaluate Multimodal Large Language Models in text and image generation tasks.

Primary intended users: The primary intended users of the Benchmark are researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence.

Contributors

BreakLee

11 commits