This repository contains guidelines and gold standard datasets for the alignment of Ancient Greek texts and translations in various languages. The guidelines are currently developed for Ancient Greek-English, Ancient Greek-Brazilian Portuguese, and Ancient Greek-Latin, and more will be added in the course of future work. Each of these guides and related gold standard was created by:
The guidelines were used to annotate a diverse dataset including Homeric epic, Attic prose, Platonic dialogue and the Fragmenta Historicorum Graecorum, and were tested by measuring inter-annotator agreement of 80% or higher. The Ancient Greek texts used are almost entirely available through the Scaife viewer (https://scaife.perseus.org/), while the fragments are from the DFHG Project (https://www.dfhg-project.org/).
The datasets used to develop the gold standard were aligned using the Ugarit Translation Alignment Editor for Historical languages (http://ugarit.ialigner.com/), and they are available in this repository.
The materials available here can be used to perform and evaluate alignments of various texts in Ancient Greek, to create new gold standard corpora, and to train automated translation models.
The guidelines can also be further adapted to address similar language pairs including an inflected and a synthetic language, such as Latin and English, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the creation of a Gold Standard in the scenario of machine translation. Different scenarios, such as language research or pedagogy, may need further tweaking to these guidelines to make them more compatible with different underlying principles.
For further information on Ugarit and translation alignment of historical languages, see http://ugarit.ialigner.com/bib.php and follow us on Twitter (@ugarit_ty).
The data in this repository is provided with a CC-BY-SA license. If you use any of the materials in this repository, please cite our work as follows.
To cite the Guidelines, please use our related publications:
If you would like to cite only one of the datasets, please use the following format:
@InProceedings{yousef-EtAl:2022:LREC,
author = {Yousef, Tariq and Palladino, Chiara and Shamsian, Farnoosh and d’Orange Ferreira, Anise and Ferreira dos Reis, Michel},
title = {An automatic model and Gold Standard for translation alignment of Ancient Greek},
booktitle = {Proceedings of the Language Resources and Evaluation Conference},
month = {June},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {5894--5905},
url = {https://aclanthology.org/2022.lrec-1.634}
}
@InProceedings{yousef-EtAl:2022:LT4HALA2022,
author = {Yousef, Tariq and Palladino, Chiara and Wright, David J. and Berti, Monica},
title = {Automatic Translation Alignment for Ancient Greek and Latin},
booktitle = {Proceedings of the Second Workshop on Language Technologies for Historical and Ancient Languages},
month = {June},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {101--107},
url = {https://aclanthology.org/2022.lt4hala2022-1.14}
}
This repository contains guidelines and gold standard datasets for the alignment of Ancient Greek texts and translations in various languages. The guidelines are currently developed for Ancient Greek-English, Ancient Greek-Brazilian Portuguese, and Ancient Greek-Latin, and more will be added in the course of future work. Each of these guides and related gold standard was created by:
The guidelines were used to annotate a diverse dataset including Homeric epic, Attic prose, Platonic dialogue and the Fragmenta Historicorum Graecorum, and were tested by measuring inter-annotator agreement of 80% or higher. The Ancient Greek texts used are almost entirely available through the Scaife viewer (https://scaife.perseus.org/), while the fragments are from the DFHG Project (https://www.dfhg-project.org/).
The datasets used to develop the gold standard were aligned using the Ugarit Translation Alignment Editor for Historical languages (http://ugarit.ialigner.com/), and they are available in this repository.
The materials available here can be used to perform and evaluate alignments of various texts in Ancient Greek, to create new gold standard corpora, and to train automated translation models.
The guidelines can also be further adapted to address similar language pairs including an inflected and a synthetic language, such as Latin and English, or can provide a structure for the alignment of other historical texts against modern translations. However, the guidelines are not project-specific: they were specifically intended for the creation of a Gold Standard in the scenario of machine translation. Different scenarios, such as language research or pedagogy, may need further tweaking to these guidelines to make them more compatible with different underlying principles.
For further information on Ugarit and translation alignment of historical languages, see http://ugarit.ialigner.com/bib.php and follow us on Twitter (@ugarit_ty).
The data in this repository is provided with a CC-BY-SA license. If you use any of the materials in this repository, please cite our work as follows.
To cite the Guidelines, please use our related publications:
If you would like to cite only one of the datasets, please use the following format:
@InProceedings{yousef-EtAl:2022:LREC,
author = {Yousef, Tariq and Palladino, Chiara and Shamsian, Farnoosh and d’Orange Ferreira, Anise and Ferreira dos Reis, Michel},
title = {An automatic model and Gold Standard for translation alignment of Ancient Greek},
booktitle = {Proceedings of the Language Resources and Evaluation Conference},
month = {June},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {5894--5905},
url = {https://aclanthology.org/2022.lrec-1.634}
}
@InProceedings{yousef-EtAl:2022:LT4HALA2022,
author = {Yousef, Tariq and Palladino, Chiara and Wright, David J. and Berti, Monica},
title = {Automatic Translation Alignment for Ancient Greek and Latin},
booktitle = {Proceedings of the Second Workshop on Language Technologies for Historical and Ancient Languages},
month = {June},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {101--107},
url = {https://aclanthology.org/2022.lt4hala2022-1.14}
}