weikaih/TaskMeAnything-v1-videoqa-2024

Dataset

Dataset Card for TaskMeAnything-v1-videoqa-2024

0

4 commits

1 linked in READMEs

updated Aug 4, 2024

See the code

README

Dataset Card for TaskMeAnything-v1-videoqa-2024

TaskMeAnything-v1-videoqa-2024 benchmark dataset

🌐 Website | πŸ“‘ Paper | πŸ€— Huggingface | πŸ’» Interface

If you like our project, please give us a star ⭐ on GitHub for latest update.

TaskMeAnything-v1-2024-Videoqa

TaskMeAnything-v1-videoqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries. This benchmark includes 2,394 3d video questions and 1,173 real video questions that the TaskMeAnything algorithm automatically approximated as challenging for over 12 popular MLMs.

The dataset contains 19 splits, while each splits contains 300+ questions from a specific task generator in TaskMeAnything-v1. For each row of dataset, it includes: video, question, options, answer and its corresponding task plan.

Load TaskMeAnything-v1-2024 VideoQA Dataset

import datasets

dataset_name = 'weikaih/TaskMeAnything-v1-videoqa-2024'
dataset = datasets.load_dataset(dataset_name, split = TASK_GENERATOR_SPLIT)

where TASK_GENERATOR_SPLIT is one of the task generators, eg, 2024_2d_how_many.

Evaluation Results

Overall

image/png

Breakdown performance on each task types

image/png

image/png

Out-of-Scope Use

This dataset should not be used for training models.

Disclaimers

TaskMeAnything and its associated resources are provided for research and educational purposes only. The authors and contributors make no warranties regarding the accuracy or reliability of the data and software. Users are responsible for ensuring their use complies with applicable laws and regulations. The project is not liable for any damages or losses resulting from the use of these resources.

Contact

Citation

BibTeX:

@article{zhang2024task,
  title={Task Me Anything},
  author={Zhang, Jieyu and Huang, Weikai and Ma, Zixian and Michel, Oscar and He, Dong and Gupta, Tanmay and Ma, Wei-Chiu and Farhadi, Ali and Kembhavi, Aniruddha and Krishna, Ranjay},
  journal={arXiv preprint arXiv:2406.11775},
  year={2024}
}

weikaih/TaskMeAnything-v1-videoqa-2024

Dataset

Dataset Card for TaskMeAnything-v1-videoqa-2024

0

4 commits

1 linked in READMEs

updated Aug 4, 2024

See the code

README

Dataset Card for TaskMeAnything-v1-videoqa-2024

TaskMeAnything-v1-videoqa-2024 benchmark dataset

🌐 Website | πŸ“‘ Paper | πŸ€— Huggingface | πŸ’» Interface

If you like our project, please give us a star ⭐ on GitHub for latest update.

TaskMeAnything-v1-2024-Videoqa

TaskMeAnything-v1-videoqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries. This benchmark includes 2,394 3d video questions and 1,173 real video questions that the TaskMeAnything algorithm automatically approximated as challenging for over 12 popular MLMs.

The dataset contains 19 splits, while each splits contains 300+ questions from a specific task generator in TaskMeAnything-v1. For each row of dataset, it includes: video, question, options, answer and its corresponding task plan.

Load TaskMeAnything-v1-2024 VideoQA Dataset

import datasets

dataset_name = 'weikaih/TaskMeAnything-v1-videoqa-2024'
dataset = datasets.load_dataset(dataset_name, split = TASK_GENERATOR_SPLIT)

where TASK_GENERATOR_SPLIT is one of the task generators, eg, 2024_2d_how_many.

Evaluation Results

Overall

image/png

Breakdown performance on each task types

image/png

image/png

Out-of-Scope Use

This dataset should not be used for training models.

Disclaimers

TaskMeAnything and its associated resources are provided for research and educational purposes only. The authors and contributors make no warranties regarding the accuracy or reliability of the data and software. Users are responsible for ensuring their use complies with applicable laws and regulations. The project is not liable for any damages or losses resulting from the use of these resources.

Contact

Citation

BibTeX:

@article{zhang2024task,
  title={Task Me Anything},
  author={Zhang, Jieyu and Huang, Weikai and Ma, Zixian and Michel, Oscar and He, Dong and Gupta, Tanmay and Ma, Wei-Chiu and Farhadi, Ali and Kembhavi, Aniruddha and Krishna, Ranjay},
  journal={arXiv preprint arXiv:2406.11775},
  year={2024}
}