luisarip/classificatio_sine_iactu

Zero-Shot NERC in Latin with GLiNER2

0

stars

16

commits

Jupyter Notebook

primary language

Mar 25, 2026

updated

README

Classificatio Sine Iactu — That Is, Zero-Shot NERC in Latin

Developed by Luisa Ripoll-Alberola1,3, Fernando Nicolás-Flores2 and Francisco Javier Muñoz Acebes3

1Computational Humanities, Leipzig University

2Departamento de Prehistoria, Arqueología, Historia Antigua, Filología Griega y Filología Latina, Universidad de Alicante

3Filología Digital, Universidad de Valladolid

Each .py file is a module containing one or more functions, intended to be imported rather than run directly as a script — for example, translator.py contains translate_latin_to_english but can be extended with additional translation backends.

  • from tsv_to_sentences import extract_sentences_from_tsv: This function reads the .tsv file properly. Rows that start with "#" are skipped, and sentences are saved separately taking the "EndOfSentence" marker at "MISC" column as the delimiter between sentences. Joining sentences this way is necessary for the step of translation and input to the zero-shot model. Some checks are included to handle cases where the first row is not parsed as a header.

  • from translator import translate_latin_to_english: Call to the Google Translator API, with some delay for the calls not to be blocked.

  • from gliner_latin import GLatiNER: Loading zero-shot NERC model (default model is fastino/gliner2-multi-v1, while passing the tagset saved in taxonomies.json. In that document, different label encodings are being saved as a dictionary. When running the model, the desired label set must be specified (in this case, "coarse_labels" or "fine_labels"). This function returns a dataframe with the original and translated sentences in parallel, with a column including entities extracted and confidence score.

  • In rule_based_processor.py one can find all rule-based correctors applied to post-process the results: in filter_entities, all entities below a confidence threshold are filtered out, and the function enforce_label_consistency ensures consistent label assignment across mentions of the same entity.

  • from alignment import align_latin, align_english: Two different functions perform the alignment, depending on whether the text has been translated. In case of working with translated English text, the alignment is performed using simalign with UGARIT/grc-alignment embeddings. Entities that are not uppercased are filtered out, and the results are converted to BIO format and projected back onto the original TSV.

There is a Jupyter Notebook containing examples of usage in the "example" folder.

Declaration on AI

During the preparation of this repository, the authors used the models ChatGPT 5, Claude Sonnet 4.5 and 4.6, Gemini 3.1 Pro, and Copilot 4, in order to draft code, debug errors, and improve code efficiency. The authors reviewed and edited the content as needed and take full responsibility for the publication’s content.

Citation

If you use this work in your research, please cite our paper:

@inproceedings{ripoll-alberola-evalatin-2026,
  author    = {Ripoll-Alberola, Luisa and Nicolás-Flores, Fernando and Muñoz Acebes, Francisco Javier},
  title     = {Classificatio Sine Iactu — That Is, Zero-Shot NERC in Latin},
  booktitle = {Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC 2026},
  year      = {2026},
  address   = {Palma de Mallorca, Spain},
  publisher = {ELRA and ICCL}
}

Contributors

luisarip

16 commits

luisarip/classificatio_sine_iactu

Zero-Shot NERC in Latin with GLiNER2

0

stars

16

commits

Jupyter Notebook

primary language

Mar 25, 2026

updated

README

Classificatio Sine Iactu — That Is, Zero-Shot NERC in Latin

Developed by Luisa Ripoll-Alberola1,3, Fernando Nicolás-Flores2 and Francisco Javier Muñoz Acebes3

1Computational Humanities, Leipzig University

2Departamento de Prehistoria, Arqueología, Historia Antigua, Filología Griega y Filología Latina, Universidad de Alicante

3Filología Digital, Universidad de Valladolid

Each .py file is a module containing one or more functions, intended to be imported rather than run directly as a script — for example, translator.py contains translate_latin_to_english but can be extended with additional translation backends.

  • from tsv_to_sentences import extract_sentences_from_tsv: This function reads the .tsv file properly. Rows that start with "#" are skipped, and sentences are saved separately taking the "EndOfSentence" marker at "MISC" column as the delimiter between sentences. Joining sentences this way is necessary for the step of translation and input to the zero-shot model. Some checks are included to handle cases where the first row is not parsed as a header.

  • from translator import translate_latin_to_english: Call to the Google Translator API, with some delay for the calls not to be blocked.

  • from gliner_latin import GLatiNER: Loading zero-shot NERC model (default model is fastino/gliner2-multi-v1, while passing the tagset saved in taxonomies.json. In that document, different label encodings are being saved as a dictionary. When running the model, the desired label set must be specified (in this case, "coarse_labels" or "fine_labels"). This function returns a dataframe with the original and translated sentences in parallel, with a column including entities extracted and confidence score.

  • In rule_based_processor.py one can find all rule-based correctors applied to post-process the results: in filter_entities, all entities below a confidence threshold are filtered out, and the function enforce_label_consistency ensures consistent label assignment across mentions of the same entity.

  • from alignment import align_latin, align_english: Two different functions perform the alignment, depending on whether the text has been translated. In case of working with translated English text, the alignment is performed using simalign with UGARIT/grc-alignment embeddings. Entities that are not uppercased are filtered out, and the results are converted to BIO format and projected back onto the original TSV.

There is a Jupyter Notebook containing examples of usage in the "example" folder.

Declaration on AI

During the preparation of this repository, the authors used the models ChatGPT 5, Claude Sonnet 4.5 and 4.6, Gemini 3.1 Pro, and Copilot 4, in order to draft code, debug errors, and improve code efficiency. The authors reviewed and edited the content as needed and take full responsibility for the publication’s content.

Citation

If you use this work in your research, please cite our paper:

@inproceedings{ripoll-alberola-evalatin-2026,
  author    = {Ripoll-Alberola, Luisa and Nicolás-Flores, Fernando and Muñoz Acebes, Francisco Javier},
  title     = {Classificatio Sine Iactu — That Is, Zero-Shot NERC in Latin},
  booktitle = {Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC 2026},
  year      = {2026},
  address   = {Palma de Mallorca, Spain},
  publisher = {ELRA and ICCL}
}

Contributors

luisarip

16 commits

Languages

Jupyter Notebook

89.1%

Python

10.9%