mlfoundations/MINT-1T

🍃 MINT-1T: A one trillion token multimodal interleaved dataset.

835

8 commits

updated Jul 31, 2024

See the code

README

MINT-1T:
Scaling Open-Source Multimodal Data by 10x:
A Multimodal Dataset with One Trillion Tokens

Paper | Dataset | Blog Post

Example Docs

🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with one trillion text tokens and 3.4 billion images, a ~10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers.

We release all subsets of MINT-1T, including:

Updates

Citation

If you found our work useful, please consider citing:

@article{awadalla2024mint1t,
      title={MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens}, 
      author={Anas Awadalla and Le Xue and Oscar Lo and Manli Shu and Hannah Lee and Etash Kumar Guha and Matt Jordan and Sheng Shen and Mohamed Awadalla and Silvio Savarese and Caiming Xiong and Ran Xu and Yejin Choi and Ludwig Schmidt},
      year={2024}
}
dataset
multimodal-learning
vision-language-model

Significant stargazers

Volodymyr Kyrylov

549 followers · starred Jun 2024

Arjen P. de Vries

106 followers · starred Jun 2024

Patrick Barker

70 followers · starred Oct 2024

Clemo

44 followers · starred Jul 2024

mlfoundations/MINT-1T

🍃 MINT-1T: A one trillion token multimodal interleaved dataset.

835

8 commits

updated Jul 31, 2024

See the code

README

MINT-1T:
Scaling Open-Source Multimodal Data by 10x:
A Multimodal Dataset with One Trillion Tokens

Paper | Dataset | Blog Post

Example Docs

🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with one trillion text tokens and 3.4 billion images, a ~10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers.

We release all subsets of MINT-1T, including:

Updates

Citation

If you found our work useful, please consider citing:

@article{awadalla2024mint1t,
      title={MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens}, 
      author={Anas Awadalla and Le Xue and Oscar Lo and Manli Shu and Hannah Lee and Etash Kumar Guha and Matt Jordan and Sheng Shen and Mohamed Awadalla and Silvio Savarese and Caiming Xiong and Ran Xu and Yejin Choi and Ludwig Schmidt},
      year={2024}
}
dataset
multimodal-learning
vision-language-model

Significant stargazers

Volodymyr Kyrylov

549 followers · starred Jun 2024

Arjen P. de Vries

106 followers · starred Jun 2024

Patrick Barker

70 followers · starred Oct 2024

Clemo

44 followers · starred Jul 2024