Jack-ZC8/M3AV-dataset

[ACL 2024] A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset

Python

24

20 commits

updated May 29, 2025

See the code

README

🎓M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset (ACL 2024)

Overview

The overview of our 🎓M3AV dataset:

  1. The first component is slides annotated with simple and complex blocks. They will be merged following some rules.
  2. The second component is speech containing special vocabulary, spoken and written forms, and word-level timestamps.
  3. The third component is the paper corresponding to the video. The asterisk (*) denotes that only computer science videos have corresponding papers.
Video FieldVideo NameTotal HoursTotal Counts
Human-computer InteractionCHI 2021 Paper Presentations (CHI)55.00660
Human-computer InteractionUbiComp 2020 Presentations (Ubi)10.65107
Biomedical SciencesNIH Director's Wednesday Afternoon Lectures (NIH)237.71228
Biomedical SciencesIntroduction to the Principles and Practice of Clinical Research (IPP)42.2767
MathematicsOxford Mathematics (MLS)27.2351

We download various academic lectures ranging from Human-computer Interaction, and Biomedical Sciences to Mathematics as shown in the table above.

News

  • [2024-08] 🤖All Benchmarks have been released!
  • [2024-06] 🤖Benchmarks of LLaMA-2 and GPT-4 have been released!
  • [2024-05] 🎉Our work has been accepted by ACL 2024 main conference!
  • [2024-04] 🔥v1.0 has been released! We have further refined all speech data. Specifically, the training set adopts the text-normalised Whisper results, and the development/testing set employs a manual combination of Whisper and microsoft STT results.

Details

The folder demo contains a sample for demonstration.

Citation

@InProceedings{Chen2024a,
  author    = {Zhe Chen and Heyang Liu and Wenyi Yu and Guangzhi Sun and Hongcheng Liu and Ji Wu and Chao Zhang and Yu Wang and Yanfeng Wang},
  booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), {ACL} 2024, Bangkok, Thailand, August 11-16, 2024},
  title     = {M{\({^3}\)}{AV}: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset},
  year      = {2024},
  editor    = {Lun{-}Wei Ku and Andre Martins and Vivek Srikumar},
  pages     = {9041--9060},
  publisher = {Association for Computational Linguistics},
  bibsource = {dblp computer science bibliography, https://dblp.org},
  biburl    = {https://dblp.org/rec/conf/acl/ChenLYSLWZWW24.bib},
  doi       = {10.18653/V1/2024.ACL-LONG.489},
}

Jack-ZC8/M3AV-dataset

[ACL 2024] A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset

Python

24

20 commits

updated May 29, 2025

See the code

README

🎓M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset (ACL 2024)

Overview

The overview of our 🎓M3AV dataset:

  1. The first component is slides annotated with simple and complex blocks. They will be merged following some rules.
  2. The second component is speech containing special vocabulary, spoken and written forms, and word-level timestamps.
  3. The third component is the paper corresponding to the video. The asterisk (*) denotes that only computer science videos have corresponding papers.
Video FieldVideo NameTotal HoursTotal Counts
Human-computer InteractionCHI 2021 Paper Presentations (CHI)55.00660
Human-computer InteractionUbiComp 2020 Presentations (Ubi)10.65107
Biomedical SciencesNIH Director's Wednesday Afternoon Lectures (NIH)237.71228
Biomedical SciencesIntroduction to the Principles and Practice of Clinical Research (IPP)42.2767
MathematicsOxford Mathematics (MLS)27.2351

We download various academic lectures ranging from Human-computer Interaction, and Biomedical Sciences to Mathematics as shown in the table above.

News

  • [2024-08] 🤖All Benchmarks have been released!
  • [2024-06] 🤖Benchmarks of LLaMA-2 and GPT-4 have been released!
  • [2024-05] 🎉Our work has been accepted by ACL 2024 main conference!
  • [2024-04] 🔥v1.0 has been released! We have further refined all speech data. Specifically, the training set adopts the text-normalised Whisper results, and the development/testing set employs a manual combination of Whisper and microsoft STT results.

Details

The folder demo contains a sample for demonstration.

Citation

@InProceedings{Chen2024a,
  author    = {Zhe Chen and Heyang Liu and Wenyi Yu and Guangzhi Sun and Hongcheng Liu and Ji Wu and Chao Zhang and Yu Wang and Yanfeng Wang},
  booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), {ACL} 2024, Bangkok, Thailand, August 11-16, 2024},
  title     = {M{\({^3}\)}{AV}: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset},
  year      = {2024},
  editor    = {Lun{-}Wei Ku and Andre Martins and Vivek Srikumar},
  pages     = {9041--9060},
  publisher = {Association for Computational Linguistics},
  bibsource = {dblp computer science bibliography, https://dblp.org},
  biburl    = {https://dblp.org/rec/conf/acl/ChenLYSLWZWW24.bib},
  doi       = {10.18653/V1/2024.ACL-LONG.489},
}

Languages

Python

53.1%

Shell

46.9%