segyges/OpenWebText2

Dataset

Dataset Card for OpenWebText2

18

6 commits

2 linked in READMEs

updated Mar 17, 2025

See the code

README

Dataset Card for OpenWebText2

OpenWebText2 is a reasonably large corpus of scraped natural language data.

Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those uses.

I am not acting on behalf of the original authors of the dataset.

More: https://openwebtext2.readthedocs.io/en/latest/

Dataset Description

  • Language(s) (NLP): English
  • License: MIT

Dataset Sources [optional]

Dataset Card Authors

SE Gyges

Dataset Card Contact

segyges on github or gmail.

segyges/OpenWebText2

Dataset

Dataset Card for OpenWebText2

18

6 commits

2 linked in READMEs

updated Mar 17, 2025

See the code

README

Dataset Card for OpenWebText2

OpenWebText2 is a reasonably large corpus of scraped natural language data.

Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those uses.

I am not acting on behalf of the original authors of the dataset.

More: https://openwebtext2.readthedocs.io/en/latest/

Dataset Description

  • Language(s) (NLP): English
  • License: MIT

Dataset Sources [optional]

Dataset Card Authors

SE Gyges

Dataset Card Contact

segyges on github or gmail.