CODAL-Bench - Evaluating LLM Alignment to Coding Preferences
6
7 commits
1 linked in READMEs
updated Mar 18, 2024
This benchmark comprises 500 random samples of CodeUltraFeedback dataset. The benchmark includes responses of multiple closed-source LLMs that can be used as references when judging other LLMs using LLM-as-a-Judge:
Please refer to our GitHub repository to evaluate your own LLM on CODAL-Bench using LLM-as-a-Judge: https://github.com/martin-wey/CodeUltraFeedback.
Please cite our work if you use our benchmark.
@article{weyssow2024codeultrafeedback,
title={CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences},
author={Weyssow, Martin and Kamanda, Aton and Sahraoui, Houari},
journal={arXiv preprint arXiv:2403.09032},
year={2024}
}
CODAL-Bench - Evaluating LLM Alignment to Coding Preferences
6
7 commits
1 linked in READMEs
updated Mar 18, 2024
This benchmark comprises 500 random samples of CodeUltraFeedback dataset. The benchmark includes responses of multiple closed-source LLMs that can be used as references when judging other LLMs using LLM-as-a-Judge:
Please refer to our GitHub repository to evaluate your own LLM on CODAL-Bench using LLM-as-a-Judge: https://github.com/martin-wey/CodeUltraFeedback.
Please cite our work if you use our benchmark.
@article{weyssow2024codeultrafeedback,
title={CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences},
author={Weyssow, Martin and Kamanda, Aton and Sahraoui, Houari},
journal={arXiv preprint arXiv:2403.09032},
year={2024}
}