harryhsing/OmniInstruct_V1_AVQA_R1

Dataset

3

stars

3

commits

1

linked in READMEs

May 8, 2025

updated

README

This repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.

For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk

Data Format in AVQA-R1-6K:

  {
      "problem_id": 0,
      "problem": "What is the source of the sound in the video?",
      "data_type": "image_audio",
      "problem_type": "multiple choice",
      "options": [
        "A. motorcycle",
        "B. automobile",
        "C. motorboat",
        "D. bus"
      ],
      "solution": "<answer>B</answer>",
      "path": {
        "image": "images/sample_0.jpg",
        "audio": "audios/sample_0.wav"
      },
      "data_source": "OmniInstruct_v1-AVQA"
  }

Contributors

harryhsing

3 commits

harryhsing/OmniInstruct_V1_AVQA_R1

Dataset

3

stars

3

commits

1

linked in READMEs

May 8, 2025

updated

README

This repository contains data presented in EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.

For training and inference, please refer to the Code: https://github.com/HarryHsing/EchoInk

Data Format in AVQA-R1-6K:

  {
      "problem_id": 0,
      "problem": "What is the source of the sound in the video?",
      "data_type": "image_audio",
      "problem_type": "multiple choice",
      "options": [
        "A. motorcycle",
        "B. automobile",
        "C. motorboat",
        "D. bus"
      ],
      "solution": "<answer>B</answer>",
      "path": {
        "image": "images/sample_0.jpg",
        "audio": "audios/sample_0.wav"
      },
      "data_source": "OmniInstruct_v1-AVQA"
  }

Contributors

harryhsing

3 commits