This repository contains the human evaluation experiment data for Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 📄. The code for the project is hosted at github.com/sarapapi/hearing2translate.
The annotations were collected using Pearmut (code), a lightweight platform that makes end-to-end human evaluation for multilingual tasks efficient and reliable.
The evaluations were done with bilingual speakers using the Pearmut tool with the MQM/ESA protocol contrastively, with multiple model outputs side by side:

Each row contains annotations for three models: aya_canary-v2, seamlessm4t, and voxtral-small-24b, in the order they were shown in columns.
The field annotations contains the score and error spans with categories.
The field src_ref contains the reference transcription of the source, though the annotators and models only heard the audio.
If you use this dataset, please cite both the research paper and the platform paper:
@misc{papi2025hearingtranslateeffectivenessspeech,
title={Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs},
author={Sara Papi and Javier Garcia Gilabert and Zachary Hopton and Vilém Zouhar and Carlos Escolano and Gerard I. Gállego and Jorge Iranzo-Sánchez and Ahrii Kim and Dominik Macháček and Patricia Schmidtova and Maike Züfle},
year={2025},
eprint={2512.16378},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2512.16378},
}
@misc{zouhar2026pearmut,
title={Pearmut: Human Evaluation of Translation Made Trivial},
author={Vilém Zouhar and Tom Kocmi},
year={2026},
eprint={2601.02933},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.02933},
}
This repository contains the human evaluation experiment data for Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 📄. The code for the project is hosted at github.com/sarapapi/hearing2translate.
The annotations were collected using Pearmut (code), a lightweight platform that makes end-to-end human evaluation for multilingual tasks efficient and reliable.
The evaluations were done with bilingual speakers using the Pearmut tool with the MQM/ESA protocol contrastively, with multiple model outputs side by side:

Each row contains annotations for three models: aya_canary-v2, seamlessm4t, and voxtral-small-24b, in the order they were shown in columns.
The field annotations contains the score and error spans with categories.
The field src_ref contains the reference transcription of the source, though the annotators and models only heard the audio.
If you use this dataset, please cite both the research paper and the platform paper:
@misc{papi2025hearingtranslateeffectivenessspeech,
title={Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs},
author={Sara Papi and Javier Garcia Gilabert and Zachary Hopton and Vilém Zouhar and Carlos Escolano and Gerard I. Gállego and Jorge Iranzo-Sánchez and Ahrii Kim and Dominik Macháček and Patricia Schmidtova and Maike Züfle},
year={2025},
eprint={2512.16378},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2512.16378},
}
@misc{zouhar2026pearmut,
title={Pearmut: Human Evaluation of Translation Made Trivial},
author={Vilém Zouhar and Tom Kocmi},
year={2026},
eprint={2601.02933},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.02933},
}