The IMPACT dataset contains 50 human created prompts for each category, 200 in total, to test LLMs general writing ability.
Instructed LLMs demonstrate promising ability in writing-based tasks, such as composing letters or ethical debates. This dataset consists prompts across 4 diverse usage scenarios:
The IMPACT dataset is included in our InstructEval Benchmark Suite.
We leverage ChatGPT to judge the quality of the generated answers by LLMs. In terms of:
Each answer is scored on a Likert scale from 1 to 5. We evaluate the models in the zero-shot setting based on the given prompt and perform sampling-based decoding with a temperature of 1.0
| Model | Size | Informative | Professional | Argumentative | Creative | Avg. | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Rel. | Coh. | Rel. | Coh. | Rel. | Coh. | Rel. | Coh. | Rel. | Coh. | ||
| ChatGPT | - | 3.34 | 3.98 | 3.88 | 3.96 | 3.96 | 3.82 | 3.92 | 3.94 | 3.78 | 3.93 |
| Flan-Alpaca | 11B | 3.56 | 3.46 | 3.54 | 3.70 | 3.22 | 3.28 | 3.70 | 3.40 | 3.51 | 3.46 |
| Dolly-V2 | 12 B | 3.54 | 3.64 | 2.96 | 3.74 | 3.66 | 3.20 | 3.02 | 3.18 | 3.30 | 3.44 |
| StableVicuna | 13B | 3.54 | 3.64 | 2.96 | 3.74 | 3.30 | 3.20 | 3.02 | 3.18 | 3.21 | 3.44 |
| Flan-T5 | 11B | 2.64 | 3.24 | 2.62 | 3.22 | 2.54 | 3.40 | 2.50 | 2.72 | 2.58 | 3.15 |
Please consider citing the following article if you found our work useful:
bibtex
@article{chia2023instructeval,
title={INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models},
author={Yew Ken Chia and Pengfei Hong and Lidong Bing and Soujanya Poria},
journal={arXiv preprint arXiv:2306.04757},
year={2023}
}
8 commits
The IMPACT dataset contains 50 human created prompts for each category, 200 in total, to test LLMs general writing ability.
Instructed LLMs demonstrate promising ability in writing-based tasks, such as composing letters or ethical debates. This dataset consists prompts across 4 diverse usage scenarios:
The IMPACT dataset is included in our InstructEval Benchmark Suite.
We leverage ChatGPT to judge the quality of the generated answers by LLMs. In terms of:
Each answer is scored on a Likert scale from 1 to 5. We evaluate the models in the zero-shot setting based on the given prompt and perform sampling-based decoding with a temperature of 1.0
| Model | Size | Informative | Professional | Argumentative | Creative | Avg. | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Rel. | Coh. | Rel. | Coh. | Rel. | Coh. | Rel. | Coh. | Rel. | Coh. | ||
| ChatGPT | - | 3.34 | 3.98 | 3.88 | 3.96 | 3.96 | 3.82 | 3.92 | 3.94 | 3.78 | 3.93 |
| Flan-Alpaca | 11B | 3.56 | 3.46 | 3.54 | 3.70 | 3.22 | 3.28 | 3.70 | 3.40 | 3.51 | 3.46 |
| Dolly-V2 | 12 B | 3.54 | 3.64 | 2.96 | 3.74 | 3.66 | 3.20 | 3.02 | 3.18 | 3.30 | 3.44 |
| StableVicuna | 13B | 3.54 | 3.64 | 2.96 | 3.74 | 3.30 | 3.20 | 3.02 | 3.18 | 3.21 | 3.44 |
| Flan-T5 | 11B | 2.64 | 3.24 | 2.62 | 3.22 | 2.54 | 3.40 | 2.50 | 2.72 | 2.58 | 3.15 |
Please consider citing the following article if you found our work useful:
bibtex
@article{chia2023instructeval,
title={INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models},
author={Yew Ken Chia and Pengfei Hong and Lidong Bing and Soujanya Poria},
journal={arXiv preprint arXiv:2306.04757},
year={2023}
}
8 commits