ZO

zouharvi/hearing2translate-humeval

Dataset

3

stars

11

commits

2

linked in READMEs

Jan 9, 2026

updated

README

This repository contains the human evaluation experiment data for Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 📄. The code for the project is hosted at github.com/sarapapi/hearing2translate.

The annotations were collected using Pearmut (code), a lightweight platform that makes end-to-end human evaluation for multilingual tasks efficient and reliable.

The evaluations were done with bilingual speakers using the Pearmut tool with the MQM/ESA protocol contrastively, with multiple model outputs side by side:

pearmut_screenshot_1

Each row contains annotations for three models: aya_canary-v2, seamlessm4t, and voxtral-small-24b, in the order they were shown in columns. The field annotations contains the score and error spans with categories. The field src_ref contains the reference transcription of the source, though the annotators and models only heard the audio.

Citations

If you use this dataset, please cite both the research paper and the platform paper:

@misc{papi2025hearingtranslateeffectivenessspeech,
      title={Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs}, 
      author={Sara Papi and Javier Garcia Gilabert and Zachary Hopton and Vilém Zouhar and Carlos Escolano and Gerard I. Gállego and Jorge Iranzo-Sánchez and Ahrii Kim and Dominik Macháček and Patricia Schmidtova and Maike Züfle},
      year={2025},
      eprint={2512.16378},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2512.16378}, 
}

@misc{zouhar2026pearmut,
      title={Pearmut: Human Evaluation of Translation Made Trivial}, 
      author={Vilém Zouhar and Tom Kocmi},
      year={2026},
      eprint={2601.02933},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2601.02933}, 
}

Contributors

VZ
Vilém Zouhar

9 commits

nielsr

1 commits

spapi

1 commits

ZO

zouharvi/hearing2translate-humeval

Dataset

3

stars

11

commits

2

linked in READMEs

Jan 9, 2026

updated

README

This repository contains the human evaluation experiment data for Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 📄. The code for the project is hosted at github.com/sarapapi/hearing2translate.

The annotations were collected using Pearmut (code), a lightweight platform that makes end-to-end human evaluation for multilingual tasks efficient and reliable.

The evaluations were done with bilingual speakers using the Pearmut tool with the MQM/ESA protocol contrastively, with multiple model outputs side by side:

pearmut_screenshot_1

Each row contains annotations for three models: aya_canary-v2, seamlessm4t, and voxtral-small-24b, in the order they were shown in columns. The field annotations contains the score and error spans with categories. The field src_ref contains the reference transcription of the source, though the annotators and models only heard the audio.

Citations

If you use this dataset, please cite both the research paper and the platform paper:

@misc{papi2025hearingtranslateeffectivenessspeech,
      title={Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs}, 
      author={Sara Papi and Javier Garcia Gilabert and Zachary Hopton and Vilém Zouhar and Carlos Escolano and Gerard I. Gállego and Jorge Iranzo-Sánchez and Ahrii Kim and Dominik Macháček and Patricia Schmidtova and Maike Züfle},
      year={2025},
      eprint={2512.16378},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2512.16378}, 
}

@misc{zouhar2026pearmut,
      title={Pearmut: Human Evaluation of Translation Made Trivial}, 
      author={Vilém Zouhar and Tom Kocmi},
      year={2026},
      eprint={2601.02933},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2601.02933}, 
}

Contributors

VZ
Vilém Zouhar

9 commits

nielsr

1 commits

spapi

1 commits