List of papers on Self-Correction of LLMs.
See the codeThis repository contains a list of papers on self-correction of large language models (LLMs) published in or before 2024.
The list is maintained by Ryo Kamoi. If you have any suggestions or corrections, please feel free to open an issue or a pull request (refer to Contributing for details).
This list is based on our survey paper. If you find this list useful, please consider citing our paper:
@article{kamoi2024self-correction,
author = {Kamoi, Ryo and Zhang, Yusen and Zhang, Nan and Han, Jiawei and Zhang, Rui},
title = "{When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs}",
journal = {Transactions of the Association for Computational Linguistics},
volume = {12},
pages = {1417-1440},
year = {2024},
month = {11},
issn = {2307-387X},
doi = {10.1162/tacl_a_00713},
url = {https://doi.org/10.1162/tacl\_a\_00713},
eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl\_a\_00713/2478635/tacl\_a\_00713.pdf},
}
Self-correction of LLMs is a framework that refines responses from LLMs using LLMs during inference. Previous work has proposed various frameworks for self-correction, such as using external tools or information, or training LLMs specifically for self-correction.
In this repository, we focus on inference-time self-correction, and differentiate it from training-time self-improvement of LLMs, which uses their own responses for improving themselves only during training.
We also do not cover generate-and-rank (or sample-and-rank), which generate multiple responses and rank them using LLMs or other models. In contrast, self-correction refines their own responses, not only selecting the best one from multiple responses.
Intrinsic self-correction is a framework that refines responses from LLMs using the same LLMs without using external feedback or training designed for self-correction.
It has been reported that intrinsic self-correction does not work well in many tasks.
Previous work has proposed self-correction frameworks that use external tools, such as code executors for code generation tasks and proof assistants for theorem proving tasks.
Previous work has proposed self-correction frameworks that use information retrieval during inference.
This section includes self-correction frameworks that train LLMs specifically for self-correction, but do not use external tools or information retrieval during inference.
It has been reported that LLMs often can self-correct their own mistakes when fine-tuned on ground-truth feedback (e.g., human annotated or generated by stronger models).
There are studies that attempt to improve self-correction without using ground-truth feedback because human annotation for self-correction is costly.
Reinforcement learning is another approach to improve self-correction without using human annotated feedback.
OpenAI o1 is a framework focusing on improving the reasoning capabilities of LLMs trained with reinforcement learning to explore multiple reasoning processes and correct their own mistakes during inference. After the release of OpenAI o1, several papers or projects have proposed frameworks similar to OpenAI o1.
For more papers related to OpenAI o1, please also refer to the following repositories.
We welcome contributions! To keep the list concise, this list will focus on papers published in or before 2024. If you’d like to add a new paper to this list, please submit a pull request. Ensure that your commit and PR have descriptive and unique titles rather than generic ones like "Updated README.md."
Kindly use the following format for your entry:
* [Short name (if exists)] **Paper Title.** *Author1, Author2, ... , and Last Author.* Conference/Journal/Preprint/Blog post. year. [[paper](url)] [[code, follow up paper, etc.](url)]
If you have any questions or suggestions, please feel free to open an issue or reach out to Ryo Kamoi (ryokamoi@psu.edu).
Please refer to LICENSE.
List of papers on Self-Correction of LLMs.
See the codeThis repository contains a list of papers on self-correction of large language models (LLMs) published in or before 2024.
The list is maintained by Ryo Kamoi. If you have any suggestions or corrections, please feel free to open an issue or a pull request (refer to Contributing for details).
This list is based on our survey paper. If you find this list useful, please consider citing our paper:
@article{kamoi2024self-correction,
author = {Kamoi, Ryo and Zhang, Yusen and Zhang, Nan and Han, Jiawei and Zhang, Rui},
title = "{When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs}",
journal = {Transactions of the Association for Computational Linguistics},
volume = {12},
pages = {1417-1440},
year = {2024},
month = {11},
issn = {2307-387X},
doi = {10.1162/tacl_a_00713},
url = {https://doi.org/10.1162/tacl\_a\_00713},
eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl\_a\_00713/2478635/tacl\_a\_00713.pdf},
}
Self-correction of LLMs is a framework that refines responses from LLMs using LLMs during inference. Previous work has proposed various frameworks for self-correction, such as using external tools or information, or training LLMs specifically for self-correction.
In this repository, we focus on inference-time self-correction, and differentiate it from training-time self-improvement of LLMs, which uses their own responses for improving themselves only during training.
We also do not cover generate-and-rank (or sample-and-rank), which generate multiple responses and rank them using LLMs or other models. In contrast, self-correction refines their own responses, not only selecting the best one from multiple responses.
Intrinsic self-correction is a framework that refines responses from LLMs using the same LLMs without using external feedback or training designed for self-correction.
It has been reported that intrinsic self-correction does not work well in many tasks.
Previous work has proposed self-correction frameworks that use external tools, such as code executors for code generation tasks and proof assistants for theorem proving tasks.
Previous work has proposed self-correction frameworks that use information retrieval during inference.
This section includes self-correction frameworks that train LLMs specifically for self-correction, but do not use external tools or information retrieval during inference.
It has been reported that LLMs often can self-correct their own mistakes when fine-tuned on ground-truth feedback (e.g., human annotated or generated by stronger models).
There are studies that attempt to improve self-correction without using ground-truth feedback because human annotation for self-correction is costly.
Reinforcement learning is another approach to improve self-correction without using human annotated feedback.
OpenAI o1 is a framework focusing on improving the reasoning capabilities of LLMs trained with reinforcement learning to explore multiple reasoning processes and correct their own mistakes during inference. After the release of OpenAI o1, several papers or projects have proposed frameworks similar to OpenAI o1.
For more papers related to OpenAI o1, please also refer to the following repositories.
We welcome contributions! To keep the list concise, this list will focus on papers published in or before 2024. If you’d like to add a new paper to this list, please submit a pull request. Ensure that your commit and PR have descriptive and unique titles rather than generic ones like "Updated README.md."
Kindly use the following format for your entry:
* [Short name (if exists)] **Paper Title.** *Author1, Author2, ... , and Last Author.* Conference/Journal/Preprint/Blog post. year. [[paper](url)] [[code, follow up paper, etc.](url)]
If you have any questions or suggestions, please feel free to open an issue or reach out to Ryo Kamoi (ryokamoi@psu.edu).
Please refer to LICENSE.