Evaluation of LLMs Should Not Ignore Non-Determinism
0
8 commits
5 linked in READMEs
updated Jul 17, 2024
Official sampling results for The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Authors: Yifan Song, Guoyin Wang, Sujian Li, Bill Yuchen Lin.
Current evaluations of large language models (LLMs) often overlook non-determinism, typically focusing on a single output per example. This limits our understanding of LLM performance variability in real-world applications. Our study addresses this issue by exploring key questions about the performance differences between greedy decoding and sampling, identifying benchmarks’ consistency regarding non-determinism, and examining unique model behaviors.
Here are our findings:
We evaluate non-determinism generation of LLMs on seven benchmarks: AlpacaEval 2, Arena-Hard, WildBench v2, MixEval, MMLU-Redux, GSM8K, and HumanEval.
| Dataset | Instance Num. | Sample Num. | Metric |
|---|---|---|---|
| AlpacaEval 2 | 805 | 16 | LC |
| Arena-Hard | 500 | 16 | Win Rate |
| WildBench v2 | 1024 | 16 | WB-Score |
| MixEval | 4000 | 16 | Score |
| MMLU-Redux | 3000 | 32 | Acc |
| GSM8K | 1319 | 128 | EM |
| HumanEval | 164 | 128 | Pass@1 |
From the results, we observe a consistent performance gap between greedy decoding and the sampling method. Greedy decoding generally proves more effective for most tasks, except for AlpacaEval.
If you find this repo helpful, please cite out paper:
@article{song2024good,
author={Yifan Song and Guoyin Wang and Sujian Li and Bill Yuchen Lin},
title={The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism},
year={2024},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
arxiv id: 2407.10457
Evaluation of LLMs Should Not Ignore Non-Determinism
0
8 commits
5 linked in READMEs
updated Jul 17, 2024
Official sampling results for The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Authors: Yifan Song, Guoyin Wang, Sujian Li, Bill Yuchen Lin.
Current evaluations of large language models (LLMs) often overlook non-determinism, typically focusing on a single output per example. This limits our understanding of LLM performance variability in real-world applications. Our study addresses this issue by exploring key questions about the performance differences between greedy decoding and sampling, identifying benchmarks’ consistency regarding non-determinism, and examining unique model behaviors.
Here are our findings:
We evaluate non-determinism generation of LLMs on seven benchmarks: AlpacaEval 2, Arena-Hard, WildBench v2, MixEval, MMLU-Redux, GSM8K, and HumanEval.
| Dataset | Instance Num. | Sample Num. | Metric |
|---|---|---|---|
| AlpacaEval 2 | 805 | 16 | LC |
| Arena-Hard | 500 | 16 | Win Rate |
| WildBench v2 | 1024 | 16 | WB-Score |
| MixEval | 4000 | 16 | Score |
| MMLU-Redux | 3000 | 32 | Acc |
| GSM8K | 1319 | 128 | EM |
| HumanEval | 164 | 128 | Pass@1 |
From the results, we observe a consistent performance gap between greedy decoding and the sampling method. Greedy decoding generally proves more effective for most tasks, except for AlpacaEval.
If you find this repo helpful, please cite out paper:
@article{song2024good,
author={Yifan Song and Guoyin Wang and Sujian Li and Bill Yuchen Lin},
title={The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism},
year={2024},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
arxiv id: 2407.10457