MIntRec2.0 is the first large-scale dataset for multimodal intent recognition and out-of-scope detection in multi-party conversations (ICLR 2024)
Python
86
26 commits
updated Aug 13, 2025
Features • Download • Dataset Description • Benchmark Framework • Quick start
MIntRec2.0 is a large-scale multimodal multi-party benchmark dataset for intent recognition and out-of-scope detection in conversations. We also provide benchmark framework and evaluation codes for usage.
Example:

| Date | Announcements |
|---|---|
| 4/2025 | The latest results of multimodal large language models on the MIntRec2.0 dataset have been released on our MMLA benchmark, with an accuracy score of over 67%. Read the paper -- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark. |
| 1/2024 | 🎆 🎆 The first large-scale multimodal intent dataset has been released. Refer to the directory MIntRec2.0 for the dataset and codes. Read the paper -- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations (Published in ICLR 2024). |
| 10/2022 | 🎆 🎆 The first multimodal intent dataset is published. Refer to the directory MIntRec for the dataset and codes. Read the paper -- MIntRec: A New Dataset for Multimodal Intent Recognition (Published in ACM MM 2022). |
MIntRec2.0 has the following features:
Large in Scale: Compared with our first version of multimodal intent recognition dataset (MIntRec), MIntRec2.0 increase the data-scale from 2.2K to 15K, with 30 intent classes, 9.3K in-scope and 5.7K out-of-scope annotated utterances with text, video, and audio modalities.
Multi-turn & Multi-party Dialogues: It contains 1,245 dialogues with an average of 12 utterances per dialogue in continuous conversations. Each utterance has an intent label in each dialogue. Each dialogue has at least two different speakers with annotated speaker identities for each utterance.
Out-of-scope Detection: As real-world dialogues are in the open-world scenarios as suggested in TEXTOIR, we further include an OOS tag for detecting those utterances that do not belong to any of existing intent classes. They can be used for out-of-distribution detection and improve system robustness.
We provide full feature (video and audio), text annotations, and raw videos of in-scope (13G) and out-of-scope data (7.44G), which can be downloaded from Google Drive.
| Item | Statistics |
|---|---|
| Number of coarse-grained intents | 2 |
| Number of fine-grained intents | 30 |
| Number of dialogues | 1,245 |
| Number of utterances | 15,040 |
| Number of words in utterances | 118,477 |
| Number of unique words in utterances | 9,524 |
| Average length of utterances | 7.0 |
| Maximum length of utterances | 46 |
| Average video clip duration | 3.0 (s) |
| Maximum video clip duration | 19.9 (s) |
| Video hours | 12.3 (h) |
We present a framework to benchmark multimodal intent understanding and out-of-scope detection in both single-turn and multi-turn conversational scenarios.
The overall framework:

The framework contains 4 main modules:
Use anaconda to create Python environment
conda create --name MIntRec python=3.9
conda activate MIntRec
Install PyTorch (Cuda version 11.2)
conda install pytorch torchvision torchaudio cudatoolkit=11.3 -c pytorch
Clone the MIntRec repository.
git clone git@github.com:thuiar/MIntRec2.0.git
cd MIntRec
Install related environmental dependencies
pip install -r requirements.txt
Run examples (Take mag-bert as an example, more can be seen here)
sh examples/run_mag_bert_baselines.sh
If this work is helpful, or you want to use the codes and results in this repo, please cite the following papers:
@inproceedings{
zhang2024mintrec,
title={{MI}ntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations},
author={Hanlei Zhang and Xin Wang and Hua Xu and Qianrui Zhou and Kai Gao and Jianhua Su and jinyue Zhao and Wenrui Li and Yanting Chen},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024},
url={https://openreview.net/forum?id=nY9nITZQjc}
}
@inproceedings{MIntRec,
author = {Zhang, Hanlei and Xu, Hua and Wang, Xin and Zhou, Qianrui and Zhao, Shaojie and Teng, Jiayan},
title = {MIntRec: A New Dataset for Multimodal Intent Recognition},
year = {2022},
booktitle = {Proceedings of the 30th ACM International Conference on Multimedia},
pages = {1688–1697},
}
The dataset and camera ready version of the paper will be updated recently.
MIntRec2.0 is the first large-scale dataset for multimodal intent recognition and out-of-scope detection in multi-party conversations (ICLR 2024)
Python
86
26 commits
updated Aug 13, 2025
Features • Download • Dataset Description • Benchmark Framework • Quick start
MIntRec2.0 is a large-scale multimodal multi-party benchmark dataset for intent recognition and out-of-scope detection in conversations. We also provide benchmark framework and evaluation codes for usage.
Example:

| Date | Announcements |
|---|---|
| 4/2025 | The latest results of multimodal large language models on the MIntRec2.0 dataset have been released on our MMLA benchmark, with an accuracy score of over 67%. Read the paper -- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark. |
| 1/2024 | 🎆 🎆 The first large-scale multimodal intent dataset has been released. Refer to the directory MIntRec2.0 for the dataset and codes. Read the paper -- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations (Published in ICLR 2024). |
| 10/2022 | 🎆 🎆 The first multimodal intent dataset is published. Refer to the directory MIntRec for the dataset and codes. Read the paper -- MIntRec: A New Dataset for Multimodal Intent Recognition (Published in ACM MM 2022). |
MIntRec2.0 has the following features:
Large in Scale: Compared with our first version of multimodal intent recognition dataset (MIntRec), MIntRec2.0 increase the data-scale from 2.2K to 15K, with 30 intent classes, 9.3K in-scope and 5.7K out-of-scope annotated utterances with text, video, and audio modalities.
Multi-turn & Multi-party Dialogues: It contains 1,245 dialogues with an average of 12 utterances per dialogue in continuous conversations. Each utterance has an intent label in each dialogue. Each dialogue has at least two different speakers with annotated speaker identities for each utterance.
Out-of-scope Detection: As real-world dialogues are in the open-world scenarios as suggested in TEXTOIR, we further include an OOS tag for detecting those utterances that do not belong to any of existing intent classes. They can be used for out-of-distribution detection and improve system robustness.
We provide full feature (video and audio), text annotations, and raw videos of in-scope (13G) and out-of-scope data (7.44G), which can be downloaded from Google Drive.
| Item | Statistics |
|---|---|
| Number of coarse-grained intents | 2 |
| Number of fine-grained intents | 30 |
| Number of dialogues | 1,245 |
| Number of utterances | 15,040 |
| Number of words in utterances | 118,477 |
| Number of unique words in utterances | 9,524 |
| Average length of utterances | 7.0 |
| Maximum length of utterances | 46 |
| Average video clip duration | 3.0 (s) |
| Maximum video clip duration | 19.9 (s) |
| Video hours | 12.3 (h) |
We present a framework to benchmark multimodal intent understanding and out-of-scope detection in both single-turn and multi-turn conversational scenarios.
The overall framework:

The framework contains 4 main modules:
Use anaconda to create Python environment
conda create --name MIntRec python=3.9
conda activate MIntRec
Install PyTorch (Cuda version 11.2)
conda install pytorch torchvision torchaudio cudatoolkit=11.3 -c pytorch
Clone the MIntRec repository.
git clone git@github.com:thuiar/MIntRec2.0.git
cd MIntRec
Install related environmental dependencies
pip install -r requirements.txt
Run examples (Take mag-bert as an example, more can be seen here)
sh examples/run_mag_bert_baselines.sh
If this work is helpful, or you want to use the codes and results in this repo, please cite the following papers:
@inproceedings{
zhang2024mintrec,
title={{MI}ntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations},
author={Hanlei Zhang and Xin Wang and Hua Xu and Qianrui Zhou and Kai Gao and Jianhua Su and jinyue Zhao and Wenrui Li and Yanting Chen},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024},
url={https://openreview.net/forum?id=nY9nITZQjc}
}
@inproceedings{MIntRec,
author = {Zhang, Hanlei and Xu, Hua and Wang, Xin and Zhou, Qianrui and Zhao, Shaojie and Teng, Jiayan},
title = {MIntRec: A New Dataset for Multimodal Intent Recognition},
year = {2022},
booktitle = {Proceedings of the 30th ACM International Conference on Multimedia},
pages = {1688–1697},
}
The dataset and camera ready version of the paper will be updated recently.