tarekeldeeb/ArabicCorpus2B

Dataset

BUILDING VOCABULARY

2

5 commits

4 linked in READMEs

updated Dec 14, 2022

See the code

README

BUILDING VOCABULARY
Processed 1754541204 tokens.
Counted 5329509 unique words.
Truncating vocabulary at min count 5.
Using vocabulary of size 1539115.

Build the Arabic Corpus

Dowload Resources

The arabic corpus {1.9B word} consists of the following resources:

  • ShamelaLibrary348.7z link {1.15B}
  • UN arabic corpus mirror1 mirror2 {0.37B}
  • AraCorpus.tar.gz link {0.14B}
  • Arabic Wikipedia Latest Articles Dump link {0.11B}
  • Tashkeela-arabic-diacritized-text-utf8-0.3.zip link {0.07B}
  • Arabic Tweets link {0.03B}
  • watan-2004.7z link {0.01B}

Build Script: https://github.com/tarekeldeeb/GloVe-Arabic/tree/master/arabic_corpus


Download the dataset

Mirror : https://archive.org/details/arabic_corpus


license: Waqf v2 (https://github.com/ojuba-org/waqf/tree/master/2.0)

tarekeldeeb/ArabicCorpus2B

Dataset

BUILDING VOCABULARY

2

5 commits

4 linked in READMEs

updated Dec 14, 2022

See the code

README

BUILDING VOCABULARY
Processed 1754541204 tokens.
Counted 5329509 unique words.
Truncating vocabulary at min count 5.
Using vocabulary of size 1539115.

Build the Arabic Corpus

Dowload Resources

The arabic corpus {1.9B word} consists of the following resources:

  • ShamelaLibrary348.7z link {1.15B}
  • UN arabic corpus mirror1 mirror2 {0.37B}
  • AraCorpus.tar.gz link {0.14B}
  • Arabic Wikipedia Latest Articles Dump link {0.11B}
  • Tashkeela-arabic-diacritized-text-utf8-0.3.zip link {0.07B}
  • Arabic Tweets link {0.03B}
  • watan-2004.7z link {0.01B}

Build Script: https://github.com/tarekeldeeb/GloVe-Arabic/tree/master/arabic_corpus


Download the dataset

Mirror : https://archive.org/details/arabic_corpus


license: Waqf v2 (https://github.com/ojuba-org/waqf/tree/master/2.0)