Nart/abkhaz_text

Dataset

4

stars

41

commits

2

linked in READMEs

Sep 17, 2024

updated

README

Dataset Card for "monolingual_ab"

Table of Contents

Dataset Description

  • Point of Contact: Nart Tlisha
  • Size of the generated dataset: 176 MB

Dataset Summary

The Abkhaz language monolingual dataset is a collection of 1,470,480 sentences extracted from different sources. The dataset is available under the Creative Commons Universal Public Domain License. Part of it is also available as part of Common Voice, another part is from the Abkhaz National Corpus

Dataset Creation

Source Data

Here is a link to the source of a large part of the data on github

Considerations for Using the Data

Other Known Limitations

The accuracy of the dataset is around 95% (gramatical, arthographical errors)

Contributors

Nart

40 commits

julien-c

1 commits

Nart/abkhaz_text

Dataset

4

stars

41

commits

2

linked in READMEs

Sep 17, 2024

updated

README

Dataset Card for "monolingual_ab"

Table of Contents

Dataset Description

  • Point of Contact: Nart Tlisha
  • Size of the generated dataset: 176 MB

Dataset Summary

The Abkhaz language monolingual dataset is a collection of 1,470,480 sentences extracted from different sources. The dataset is available under the Creative Commons Universal Public Domain License. Part of it is also available as part of Common Voice, another part is from the Abkhaz National Corpus

Dataset Creation

Source Data

Here is a link to the source of a large part of the data on github

Considerations for Using the Data

Other Known Limitations

The accuracy of the dataset is around 95% (gramatical, arthographical errors)

Contributors

Nart

40 commits

julien-c

1 commits