ColorfulAI/OmniMMI

Dataset

OmniMMI

0

7 commits

2 linked in READMEs

updated Apr 3, 2025

See the code

README

OmniMMI

Paper: OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

Code

Dataset Description

we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 interactive videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six distinct subtasks.

  • Streaming Temporal State Awareness. Streaming video understanding must build an understanding w.r.t. the current and historical temporal state incrementally, without accessing the future context. This contrasts with traditional MLLM that can leverage the entire multi-modal contexts, posing challenges in our distinguished tasks of action prediction (AP), state grounding (SG) and multi-turn dependencies (MD).

  • Proactive Reasoning and Turn-Taking. Generating responses proactively and appropriately anticipating the turn-taking time spot w.r.t. user's intentions and dynamic contexts is a crucial feature for general interactive agents. This typically requires models to identify speakers (SI), distinguish between noise or legitimate query (PT), and proactively initiate a response (PA).

images

Data Statistics

StatisticSGAPMPPTPASI
Videos30020030078200200
Queries704200786200200200
Avg. Turns2.351.002.621.001.001.00
Avg. Vid.(s)350.82234.95374.802004.10149.82549.64
Avg. Que.16.0025.9926.278.4517.4960.91

Evaluation

We provide OmniMMI for evaluation

Leaderboard

See our page

Point of Contact: mailto:Yuxuan Wang

Contributors

ColorfulAI

6 commits

nielsr

1 commits

ColorfulAI/OmniMMI

Dataset

OmniMMI

0

7 commits

2 linked in READMEs

updated Apr 3, 2025

See the code

README

OmniMMI

Paper: OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

Code

Dataset Description

we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 interactive videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six distinct subtasks.

  • Streaming Temporal State Awareness. Streaming video understanding must build an understanding w.r.t. the current and historical temporal state incrementally, without accessing the future context. This contrasts with traditional MLLM that can leverage the entire multi-modal contexts, posing challenges in our distinguished tasks of action prediction (AP), state grounding (SG) and multi-turn dependencies (MD).

  • Proactive Reasoning and Turn-Taking. Generating responses proactively and appropriately anticipating the turn-taking time spot w.r.t. user's intentions and dynamic contexts is a crucial feature for general interactive agents. This typically requires models to identify speakers (SI), distinguish between noise or legitimate query (PT), and proactively initiate a response (PA).

images

Data Statistics

StatisticSGAPMPPTPASI
Videos30020030078200200
Queries704200786200200200
Avg. Turns2.351.002.621.001.001.00
Avg. Vid.(s)350.82234.95374.802004.10149.82549.64
Avg. Que.16.0025.9926.278.4517.4960.91

Evaluation

We provide OmniMMI for evaluation

Leaderboard

See our page

Point of Contact: mailto:Yuxuan Wang

Contributors

ColorfulAI

6 commits

nielsr

1 commits