AmourWaltz/Awesome-Reliable-LLM

JavaScript

200

41 commits

updated Sep 21, 2026

See the code

README

Awesome Reliable LLMs

Factuality · Honesty · Consistency

Three connected motifs for reliable LLMs: verified evidence, calibrated confidence, and coherent context.

A curated literature collection on building reliable large language models to mitigate hallucination.

287 curated papers 3 dimensions and 13 topics Bibliographic information checked on 2026-09-22

Background  ·  Factuality  ·  Honesty  ·  Consistency  ·  Contribute


This project follows the research themes of my PhD thesis, A Framework of Building Reliable Large Language Models to Mitigate Hallucination, and curates related literature on factuality, honesty, and consistency.

Introduction

Mitigating hallucination requires understanding both why models produce unreliable responses and how to improve their behavior. Reliability encompasses factual accuracy, awareness of knowledge limitations, and coherence with relevant context.

A reliable LLM should use knowledge to produce factually accurate responses, honestly communicate uncertainty and knowledge limitations, and maintain contextual consistency throughout multi-turn interactions.

DimensionCentral questionLiterature focus
FactualityDoes the response agree with verifiable facts?Factual knowledge, knowledge boundaries, hallucination detection, and factuality improvement.
HonestyDoes the model accurately communicate what it knows and does not know?Self-awareness, confidence estimation, calibration, and uncertainty expression.
ConsistencyDoes the response remain coherent with relevant evidence and the interaction history?Contextual faithfulness, multi-turn dialogue, retrieval, and search agents.

Research on hallucination, knowledge, and uncertainty is organized within these three dimensions. Papers are grouped by their primary research question, with cross-references for work spanning multiple dimensions.

Outline

Browse all sections and research topics
How to read the bibliography

Papers are grouped by their primary research focus and ordered by year, newest first. Each entry links to the paper and gives its conference, workshop, or journal and year. Papers with a verified acceptance but no linked proceedings version are marked accepted. Where a formal venue has not been verified, the entry is labeled arXiv preprint with its initial submission month. Bibliographic information was checked on 2026-09-22.


Motivation and Background

Hallucination: Definitions and Scope

Large language models have advanced from text generation to agents that retrieve information, use tools, and interact with external environments. Yet they can still produce factually incorrect, fabricated, deceptive, or inconsistent content. These hallucination-related failures pose a central challenge to LLM reliability.

Such failures can mislead users even when the responses appear fluent and confident. The examples below illustrate this concern in health care and economics, motivating the need for reliable generation in knowledge-intensive applications.

Two illustrative hallucinated responses in health care and economics, with problematic claims highlighted in red.

Illustrative hallucinated responses in health care and economics; highlighted claims demonstrate unreliable outputs.

Causes of Hallucination

A central challenge is the gap between learning frequent textual patterns and reliably applying factual knowledge. During pre-training, next-token prediction teaches models to produce plausible continuations from large text corpora. During post-training, models learn response patterns from instruction and feedback data. Learning these patterns does not by itself ensure factual understanding or a clear boundary between known and unknown information.

The example below illustrates how a familiar statement about the first person on the Moon can be reused to answer a superficially similar question about Mars. This pattern-matching perspective connects the training process to the unreliable behaviors discussed next.

A model transfers a familiar training statement about the Moon to a question about Mars, illustrating an unsupported answer from pattern matching.

An illustration of hallucination arising from reliance on familiar textual patterns.

Three Manifestations of Unreliability

Unreliable behavior manifests at three levels: the model's expression of its own knowledge, the factual content of a single response, and the coherence of responses across an interaction.

  1. Dishonesty: the model claims knowledge or expresses unwarranted certainty when it should acknowledge uncertainty or limitations.
  2. Non-factual responses: an individual answer contains incorrect or unsupported factual claims.
  3. Contextual inconsistency: responses contradict earlier statements or depart from relevant context during multi-turn interactions.

Three manifestations of unreliability: dishonesty about knowledge, non-factual content in a single answer, and inconsistency across multiple turns.

Three manifestations of unreliability: dishonesty, non-factual responses, and contextual inconsistency.

These failure modes motivate complementary reliability requirements. Assessing answer correctness alone does not capture whether a model communicates its limitations appropriately or remains consistent throughout a conversation.

From Hallucination Mitigation to LLM Reliability

Research on hallucination mitigation spans data enhancement, fine-tuning, and retrieval-augmented generation (RAG). The causes and manifestations of hallucination motivate three complementary objectives for reliable LLMs:

  • Factuality: use available knowledge to generate factually accurate responses.
  • Honesty: communicate uncertainty and acknowledge knowledge limitations when faced with unsure queries.
  • Consistency: maintain coherence with relevant context throughout multi-turn interactions.

A conceptual framework connecting textual pattern learning to unreliable outputs and motivating factuality, honesty, and consistency as complementary objectives.

A conceptual framework connecting hallucination causes, unreliable outputs, and the three dimensions of LLM reliability.

Following this progression, the literature collection is organized into Factuality, Honesty, and Consistency. Each section covers the corresponding research questions, related approaches, and evaluation settings, with links across dimensions where their concerns overlap.

↑ Back to top


Factuality

100 papers across 5 topics

Focus: the factual correctness of generated content, especially in knowledge-intensive responses.

Improving factuality involves understanding knowledge boundaries, detecting factual errors, improving responses through alignment and inference, and evaluating the resulting generations. The following reading lists cover these research questions, methods, and evaluation resources.

Knowledge Boundary

Studies of what a model knows, how reliably it can access that knowledge, and how prompting, retrieval, and fine-tuning affect its knowledge limits.

Hallucination Detection

Methods for identifying factual errors or estimating hallucination risk using sampled responses, semantic uncertainty, internal states, and reasoning behavior. These approaches connect hallucination detection with probing a model's knowledge and uncertainty.

Factuality Alignment

Training approaches that improve factual responses through data construction, supervised fine-tuning, factuality preferences, reinforcement learning, or knowledge distillation. This category also includes learning to refuse questions beyond the model's knowledge.

Factuality Inference

Approaches that improve factual generation through retrieval, prompting, sampling, decoding, activation intervention, or verification and correction. Retrieval systems that also require training are included here for their generation-time use of external evidence.

Factuality Evaluation

Benchmarks and evaluation methods for factual correctness, hallucination recognition, and factual support in generated text. The list includes question-answering datasets such as TriviaQA, SciQ, and Natural Questions, together with the work introducing the open-domain NQ setup. Coverage extends to short-form answers, long-form claims, and retrieval-grounded generation.

↑ Back to top


Honesty

94 papers across 5 topics

Focus: recognizing knowledge and capability limits, and faithfully communicating uncertainty about generated responses.

Honesty involves recognizing knowledge and capability limits, estimating and calibrating confidence, and appropriately expressing uncertainty. It is model-specific: evaluating honesty requires assessing how a model's expressed certainty and answering behavior relate to its own knowledge and capabilities.

Self-Awareness and Knowledge Limits

Studies of a model's ability to assess what it knows, recognize insufficient information, false premises, or ill-posed problems, and decide when to seek external evidence. This includes identifying mathematical unsolvability and distinguishing it from the model's own capability limits. Related studies of known and unknown questions are also collected under Knowledge Boundary.

Confidence and Uncertainty Estimation

Methods for estimating uncertainty from token probabilities, verbalized confidence, sampled responses, semantic variation, and internal representations. The list also includes uncertainty signals used to guide learning and agent decisions. Applications to factual-error detection are collected under Hallucination Detection.

Overview of likelihood-based, prompting-based, sampling-based, and training-based confidence and uncertainty estimation methods and their limitations

Overview of four families of confidence and uncertainty estimation methods—likelihood-based, prompting-based, sampling-based, and training-based—and their limitations.

Confidence Calibration

Methods for aligning confidence estimates with observed correctness, including prompting, post-hoc recalibration, supervised learning, and reinforcement learning. Coverage extends from calibration foundations to long-form generation, confidence in reasoning steps and final answers, multi-turn conversations, and tool-using agents.

Uncertainty Expression and Abstention

Research on communicating uncertainty in words or numbers, explaining uncertainty, asking for clarification, and appropriately declining to answer. The goal is to acknowledge limitations while retaining useful answers to answerable questions. Related refusal-training approaches, including R-Tuning and RL from knowledge feedback, appear under Factuality Alignment.

Honesty Evaluation

Surveys, benchmarks, and evaluation studies for self-knowledge, confidence estimation, calibration, and faithful uncertainty expression. Evaluation considers discrimination between correct and incorrect answers, calibration, appropriate abstention, and robustness across tasks and interaction settings. Evaluation settings include mathematical solvability, theorem-proving sycophancy, and multilingual confidence estimation.

↑ Back to top


Consistency

93 papers across 3 topics

Focus: agreement with relevant evidence, task constraints, and interaction history, including multi-turn dialogue, RAG, and search agents.

Maintaining consistency requires understanding contextual faithfulness and the effects of context interference, developing mitigation methods, and evaluating behavior across interactions. When sources conflict or new evidence becomes available, it also involves resolving conflicts and appropriately updating earlier claims.

Contextual Consistency and Faithfulness

Studies of how generated content relates to source material, task instructions, and interaction history. Coverage includes context–memory conflicts, conflicts among external sources, irrelevant-context interference, long-context information loss, and contradictions across conversation turns. These studies connect source faithfulness in summarization and dialogue to reliability in multi-turn agents.

Improving Contextual Consistency

Approaches for improving source faithfulness and coherence across interactions. Coverage includes key-information extraction and reranking, context compression, and prompting, together with training for context adherence, conflict-aware decoding, verification, and response correction. IRCoT, ReAct, Search-o1, Search-R1, and RAG-Gym extend this coverage to iterative retrieval, reasoning, and agent training. Related work on context-aware decoding, Self-RAG, and general retrieval-augmented generation is listed under Factuality Inference.

Consistency Evaluation

Benchmarks, metrics, and human-evaluation protocols for source support, contradiction detection, knowledge-conflict handling, conversational memory, and consistency across turns. Coverage spans summarization, data-to-text generation, knowledge-grounded dialogue, RAG, and agent interactions. Question-answering datasets provide task-level evaluation resources, complemented by checks of source support and consistency across turns. Evaluation considers evidence use and context adherence alongside answer correctness, task success, and interaction cost. Related resources, including Natural Questions, TriviaQA, RAGTruth, DeepTRACE, and long-document factual-consistency stress tests, are listed under Factuality Evaluation.

↑ Back to top


Open Challenges and Future Directions

  • Joint evaluation and improvement of factuality, honesty, and consistency.
  • Reliable long-form generation and extended agent interactions.
  • Generalization across languages, domains, and reasoning tasks.
  • Reliability in multimodal LLMs and agents.
  • Balancing factual accuracy, appropriate abstention, contextual consistency, and computational cost.

Contributing

Contributions of relevant papers, surveys, benchmarks, and tools are welcome. Place each work under the subsection matching its main research question and cross-reference other dimensions when useful.

Use a linked title followed by the publication venue and year, matching the existing reading lists:

- [Paper title](PAPER_URL) — **Venue YYYY**.

Verify bibliographic details before adding an entry. Briefly describe its relevance when helpful, and include code or data links when available.

↑ Back to top

hallucination
knowledge
reliable
uncertainty

AmourWaltz/Awesome-Reliable-LLM

JavaScript

200

41 commits

updated Sep 21, 2026

See the code

README

Awesome Reliable LLMs

Factuality · Honesty · Consistency

Three connected motifs for reliable LLMs: verified evidence, calibrated confidence, and coherent context.

A curated literature collection on building reliable large language models to mitigate hallucination.

287 curated papers 3 dimensions and 13 topics Bibliographic information checked on 2026-09-22

Background  ·  Factuality  ·  Honesty  ·  Consistency  ·  Contribute


This project follows the research themes of my PhD thesis, A Framework of Building Reliable Large Language Models to Mitigate Hallucination, and curates related literature on factuality, honesty, and consistency.

Introduction

Mitigating hallucination requires understanding both why models produce unreliable responses and how to improve their behavior. Reliability encompasses factual accuracy, awareness of knowledge limitations, and coherence with relevant context.

A reliable LLM should use knowledge to produce factually accurate responses, honestly communicate uncertainty and knowledge limitations, and maintain contextual consistency throughout multi-turn interactions.

DimensionCentral questionLiterature focus
FactualityDoes the response agree with verifiable facts?Factual knowledge, knowledge boundaries, hallucination detection, and factuality improvement.
HonestyDoes the model accurately communicate what it knows and does not know?Self-awareness, confidence estimation, calibration, and uncertainty expression.
ConsistencyDoes the response remain coherent with relevant evidence and the interaction history?Contextual faithfulness, multi-turn dialogue, retrieval, and search agents.

Research on hallucination, knowledge, and uncertainty is organized within these three dimensions. Papers are grouped by their primary research question, with cross-references for work spanning multiple dimensions.

Outline

Browse all sections and research topics
How to read the bibliography

Papers are grouped by their primary research focus and ordered by year, newest first. Each entry links to the paper and gives its conference, workshop, or journal and year. Papers with a verified acceptance but no linked proceedings version are marked accepted. Where a formal venue has not been verified, the entry is labeled arXiv preprint with its initial submission month. Bibliographic information was checked on 2026-09-22.


Motivation and Background

Hallucination: Definitions and Scope

Large language models have advanced from text generation to agents that retrieve information, use tools, and interact with external environments. Yet they can still produce factually incorrect, fabricated, deceptive, or inconsistent content. These hallucination-related failures pose a central challenge to LLM reliability.

Such failures can mislead users even when the responses appear fluent and confident. The examples below illustrate this concern in health care and economics, motivating the need for reliable generation in knowledge-intensive applications.

Two illustrative hallucinated responses in health care and economics, with problematic claims highlighted in red.

Illustrative hallucinated responses in health care and economics; highlighted claims demonstrate unreliable outputs.

Causes of Hallucination

A central challenge is the gap between learning frequent textual patterns and reliably applying factual knowledge. During pre-training, next-token prediction teaches models to produce plausible continuations from large text corpora. During post-training, models learn response patterns from instruction and feedback data. Learning these patterns does not by itself ensure factual understanding or a clear boundary between known and unknown information.

The example below illustrates how a familiar statement about the first person on the Moon can be reused to answer a superficially similar question about Mars. This pattern-matching perspective connects the training process to the unreliable behaviors discussed next.

A model transfers a familiar training statement about the Moon to a question about Mars, illustrating an unsupported answer from pattern matching.

An illustration of hallucination arising from reliance on familiar textual patterns.

Three Manifestations of Unreliability

Unreliable behavior manifests at three levels: the model's expression of its own knowledge, the factual content of a single response, and the coherence of responses across an interaction.

  1. Dishonesty: the model claims knowledge or expresses unwarranted certainty when it should acknowledge uncertainty or limitations.
  2. Non-factual responses: an individual answer contains incorrect or unsupported factual claims.
  3. Contextual inconsistency: responses contradict earlier statements or depart from relevant context during multi-turn interactions.

Three manifestations of unreliability: dishonesty about knowledge, non-factual content in a single answer, and inconsistency across multiple turns.

Three manifestations of unreliability: dishonesty, non-factual responses, and contextual inconsistency.

These failure modes motivate complementary reliability requirements. Assessing answer correctness alone does not capture whether a model communicates its limitations appropriately or remains consistent throughout a conversation.

From Hallucination Mitigation to LLM Reliability

Research on hallucination mitigation spans data enhancement, fine-tuning, and retrieval-augmented generation (RAG). The causes and manifestations of hallucination motivate three complementary objectives for reliable LLMs:

  • Factuality: use available knowledge to generate factually accurate responses.
  • Honesty: communicate uncertainty and acknowledge knowledge limitations when faced with unsure queries.
  • Consistency: maintain coherence with relevant context throughout multi-turn interactions.

A conceptual framework connecting textual pattern learning to unreliable outputs and motivating factuality, honesty, and consistency as complementary objectives.

A conceptual framework connecting hallucination causes, unreliable outputs, and the three dimensions of LLM reliability.

Following this progression, the literature collection is organized into Factuality, Honesty, and Consistency. Each section covers the corresponding research questions, related approaches, and evaluation settings, with links across dimensions where their concerns overlap.

↑ Back to top


Factuality

100 papers across 5 topics

Focus: the factual correctness of generated content, especially in knowledge-intensive responses.

Improving factuality involves understanding knowledge boundaries, detecting factual errors, improving responses through alignment and inference, and evaluating the resulting generations. The following reading lists cover these research questions, methods, and evaluation resources.

Knowledge Boundary

Studies of what a model knows, how reliably it can access that knowledge, and how prompting, retrieval, and fine-tuning affect its knowledge limits.

Hallucination Detection

Methods for identifying factual errors or estimating hallucination risk using sampled responses, semantic uncertainty, internal states, and reasoning behavior. These approaches connect hallucination detection with probing a model's knowledge and uncertainty.

Factuality Alignment

Training approaches that improve factual responses through data construction, supervised fine-tuning, factuality preferences, reinforcement learning, or knowledge distillation. This category also includes learning to refuse questions beyond the model's knowledge.

Factuality Inference

Approaches that improve factual generation through retrieval, prompting, sampling, decoding, activation intervention, or verification and correction. Retrieval systems that also require training are included here for their generation-time use of external evidence.

Factuality Evaluation

Benchmarks and evaluation methods for factual correctness, hallucination recognition, and factual support in generated text. The list includes question-answering datasets such as TriviaQA, SciQ, and Natural Questions, together with the work introducing the open-domain NQ setup. Coverage extends to short-form answers, long-form claims, and retrieval-grounded generation.

↑ Back to top


Honesty

94 papers across 5 topics

Focus: recognizing knowledge and capability limits, and faithfully communicating uncertainty about generated responses.

Honesty involves recognizing knowledge and capability limits, estimating and calibrating confidence, and appropriately expressing uncertainty. It is model-specific: evaluating honesty requires assessing how a model's expressed certainty and answering behavior relate to its own knowledge and capabilities.

Self-Awareness and Knowledge Limits

Studies of a model's ability to assess what it knows, recognize insufficient information, false premises, or ill-posed problems, and decide when to seek external evidence. This includes identifying mathematical unsolvability and distinguishing it from the model's own capability limits. Related studies of known and unknown questions are also collected under Knowledge Boundary.

Confidence and Uncertainty Estimation

Methods for estimating uncertainty from token probabilities, verbalized confidence, sampled responses, semantic variation, and internal representations. The list also includes uncertainty signals used to guide learning and agent decisions. Applications to factual-error detection are collected under Hallucination Detection.

Overview of likelihood-based, prompting-based, sampling-based, and training-based confidence and uncertainty estimation methods and their limitations

Overview of four families of confidence and uncertainty estimation methods—likelihood-based, prompting-based, sampling-based, and training-based—and their limitations.

Confidence Calibration

Methods for aligning confidence estimates with observed correctness, including prompting, post-hoc recalibration, supervised learning, and reinforcement learning. Coverage extends from calibration foundations to long-form generation, confidence in reasoning steps and final answers, multi-turn conversations, and tool-using agents.

Uncertainty Expression and Abstention

Research on communicating uncertainty in words or numbers, explaining uncertainty, asking for clarification, and appropriately declining to answer. The goal is to acknowledge limitations while retaining useful answers to answerable questions. Related refusal-training approaches, including R-Tuning and RL from knowledge feedback, appear under Factuality Alignment.

Honesty Evaluation

Surveys, benchmarks, and evaluation studies for self-knowledge, confidence estimation, calibration, and faithful uncertainty expression. Evaluation considers discrimination between correct and incorrect answers, calibration, appropriate abstention, and robustness across tasks and interaction settings. Evaluation settings include mathematical solvability, theorem-proving sycophancy, and multilingual confidence estimation.

↑ Back to top


Consistency

93 papers across 3 topics

Focus: agreement with relevant evidence, task constraints, and interaction history, including multi-turn dialogue, RAG, and search agents.

Maintaining consistency requires understanding contextual faithfulness and the effects of context interference, developing mitigation methods, and evaluating behavior across interactions. When sources conflict or new evidence becomes available, it also involves resolving conflicts and appropriately updating earlier claims.

Contextual Consistency and Faithfulness

Studies of how generated content relates to source material, task instructions, and interaction history. Coverage includes context–memory conflicts, conflicts among external sources, irrelevant-context interference, long-context information loss, and contradictions across conversation turns. These studies connect source faithfulness in summarization and dialogue to reliability in multi-turn agents.

Improving Contextual Consistency

Approaches for improving source faithfulness and coherence across interactions. Coverage includes key-information extraction and reranking, context compression, and prompting, together with training for context adherence, conflict-aware decoding, verification, and response correction. IRCoT, ReAct, Search-o1, Search-R1, and RAG-Gym extend this coverage to iterative retrieval, reasoning, and agent training. Related work on context-aware decoding, Self-RAG, and general retrieval-augmented generation is listed under Factuality Inference.

Consistency Evaluation

Benchmarks, metrics, and human-evaluation protocols for source support, contradiction detection, knowledge-conflict handling, conversational memory, and consistency across turns. Coverage spans summarization, data-to-text generation, knowledge-grounded dialogue, RAG, and agent interactions. Question-answering datasets provide task-level evaluation resources, complemented by checks of source support and consistency across turns. Evaluation considers evidence use and context adherence alongside answer correctness, task success, and interaction cost. Related resources, including Natural Questions, TriviaQA, RAGTruth, DeepTRACE, and long-document factual-consistency stress tests, are listed under Factuality Evaluation.

↑ Back to top


Open Challenges and Future Directions

  • Joint evaluation and improvement of factuality, honesty, and consistency.
  • Reliable long-form generation and extended agent interactions.
  • Generalization across languages, domains, and reasoning tasks.
  • Reliability in multimodal LLMs and agents.
  • Balancing factual accuracy, appropriate abstention, contextual consistency, and computational cost.

Contributing

Contributions of relevant papers, surveys, benchmarks, and tools are welcome. Place each work under the subsection matching its main research question and cross-reference other dimensions when useful.

Use a linked title followed by the publication venue and year, matching the existing reading lists:

- [Paper title](PAPER_URL) — **Venue YYYY**.

Verify bibliographic details before adding an entry. Briefly describe its relevance when helpful, and include code or data links when available.

↑ Back to top

hallucination
knowledge
reliable
uncertainty

Languages

JavaScript

96.9%

CSS

2.2%