facebook/menlo

Dataset

This dataset is released as part of MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 languages.

13

6 commits

1 linked in READMEs

updated Oct 1, 2025

See the code

README

description

This dataset is released as part of MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 languages.

MENLO

tl;dr: Massively multilingual preference evaluation, reward modeling, and post-training to improve LLMs' language proficiency

Ensuring native-like quality of large language model (LLM) responses across many languages is challenging. To address this, we introduce MENLO, a framework that operationalizes the evaluation of native-like response quality based on audience design-inspired mechanisms. Using MENLO, we create a dataset of 6,423 human-annotated prompt–response preference pairs covering four quality dimensions with high inter-annotator agreement in 47 language varieties. Our evaluation reveals that zero-shot LLM judges benefit significantly from pairwise evaluation and our structured annotation rubrics, yet they still underperform human annotators on our dataset. We demonstrate substantial improvements through fine-tuning with reinforcement learning, reward shaping, and multi-task learning approaches. Additionally, we show that RL-trained judges can serve as generative reward models to enhance LLMs' multilingual proficiency, though discrepancies with human judgment remain. Our findings suggest promising directions for scalable multilingual evaluation and preference alignment. We release our dataset and evaluation framework to support further research in multilingual LLM evaluation.

For more details, please refer to our MENLO paper.

Citation

If you use the MENLO dataset from our work, please cite with the following BibTex entry:

@article{whitehouse2025menlo,
      title={MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages}, 
      author={Chenxi Whitehouse and Sebastian Ruder and Tony Lin and Oksana Kurylo and Haruka Takagi and Janice Lam and Nicolò Busetto and Denise Diaz},
      year={2025},
      journal={arXiv preprint arXiv:2509.26601},
      url={https://arxiv.org/abs/2509.26601}, 
}

License

Use of this repository and related resources are governed by MENLO Research License.

human_evaluation
multilingual
reward_model

Contributors

chenxwh-meta

5 commits

meta-bot

1 commits

facebook/menlo

Dataset

This dataset is released as part of MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 languages.

13

6 commits

1 linked in READMEs

updated Oct 1, 2025

See the code

README

description

This dataset is released as part of MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 languages.

MENLO

tl;dr: Massively multilingual preference evaluation, reward modeling, and post-training to improve LLMs' language proficiency

Ensuring native-like quality of large language model (LLM) responses across many languages is challenging. To address this, we introduce MENLO, a framework that operationalizes the evaluation of native-like response quality based on audience design-inspired mechanisms. Using MENLO, we create a dataset of 6,423 human-annotated prompt–response preference pairs covering four quality dimensions with high inter-annotator agreement in 47 language varieties. Our evaluation reveals that zero-shot LLM judges benefit significantly from pairwise evaluation and our structured annotation rubrics, yet they still underperform human annotators on our dataset. We demonstrate substantial improvements through fine-tuning with reinforcement learning, reward shaping, and multi-task learning approaches. Additionally, we show that RL-trained judges can serve as generative reward models to enhance LLMs' multilingual proficiency, though discrepancies with human judgment remain. Our findings suggest promising directions for scalable multilingual evaluation and preference alignment. We release our dataset and evaluation framework to support further research in multilingual LLM evaluation.

For more details, please refer to our MENLO paper.

Citation

If you use the MENLO dataset from our work, please cite with the following BibTex entry:

@article{whitehouse2025menlo,
      title={MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages}, 
      author={Chenxi Whitehouse and Sebastian Ruder and Tony Lin and Oksana Kurylo and Haruka Takagi and Janice Lam and Nicolò Busetto and Denise Diaz},
      year={2025},
      journal={arXiv preprint arXiv:2509.26601},
      url={https://arxiv.org/abs/2509.26601}, 
}

License

Use of this repository and related resources are governed by MENLO Research License.

human_evaluation
multilingual
reward_model

Contributors

chenxwh-meta

5 commits

meta-bot

1 commits