Factuality · Honesty · Consistency
A curated literature collection on building reliable large language models to mitigate hallucination.
Background · Factuality · Honesty · Consistency · Contribute
This project follows the research themes of my PhD thesis, A Framework of Building Reliable Large Language Models to Mitigate Hallucination, and curates related literature on factuality, honesty, and consistency.
Mitigating hallucination requires understanding both why models produce unreliable responses and how to improve their behavior. Reliability encompasses factual accuracy, awareness of knowledge limitations, and coherence with relevant context.
A reliable LLM should use knowledge to produce factually accurate responses, honestly communicate uncertainty and knowledge limitations, and maintain contextual consistency throughout multi-turn interactions.
| Dimension | Central question | Literature focus |
|---|---|---|
| Factuality | Does the response agree with verifiable facts? | Factual knowledge, knowledge boundaries, hallucination detection, and factuality improvement. |
| Honesty | Does the model accurately communicate what it knows and does not know? | Self-awareness, confidence estimation, calibration, and uncertainty expression. |
| Consistency | Does the response remain coherent with relevant evidence and the interaction history? | Contextual faithfulness, multi-turn dialogue, retrieval, and search agents. |
Research on hallucination, knowledge, and uncertainty is organized within these three dimensions. Papers are grouped by their primary research question, with cross-references for work spanning multiple dimensions.
Papers are grouped by their primary research focus and ordered by year, newest first. Each entry links to the paper and gives its conference, workshop, or journal and year. Papers with a verified acceptance but no linked proceedings version are marked accepted. Where a formal venue has not been verified, the entry is labeled arXiv preprint with its initial submission month. Bibliographic information was checked on 2026-09-22.
Large language models have advanced from text generation to agents that retrieve information, use tools, and interact with external environments. Yet they can still produce factually incorrect, fabricated, deceptive, or inconsistent content. These hallucination-related failures pose a central challenge to LLM reliability.
Such failures can mislead users even when the responses appear fluent and confident. The examples below illustrate this concern in health care and economics, motivating the need for reliable generation in knowledge-intensive applications.
Illustrative hallucinated responses in health care and economics; highlighted claims demonstrate unreliable outputs.
A central challenge is the gap between learning frequent textual patterns and reliably applying factual knowledge. During pre-training, next-token prediction teaches models to produce plausible continuations from large text corpora. During post-training, models learn response patterns from instruction and feedback data. Learning these patterns does not by itself ensure factual understanding or a clear boundary between known and unknown information.
The example below illustrates how a familiar statement about the first person on the Moon can be reused to answer a superficially similar question about Mars. This pattern-matching perspective connects the training process to the unreliable behaviors discussed next.
An illustration of hallucination arising from reliance on familiar textual patterns.
Unreliable behavior manifests at three levels: the model's expression of its own knowledge, the factual content of a single response, and the coherence of responses across an interaction.
Three manifestations of unreliability: dishonesty, non-factual responses, and contextual inconsistency.
These failure modes motivate complementary reliability requirements. Assessing answer correctness alone does not capture whether a model communicates its limitations appropriately or remains consistent throughout a conversation.
Research on hallucination mitigation spans data enhancement, fine-tuning, and retrieval-augmented generation (RAG). The causes and manifestations of hallucination motivate three complementary objectives for reliable LLMs:
A conceptual framework connecting hallucination causes, unreliable outputs, and the three dimensions of LLM reliability.
Following this progression, the literature collection is organized into Factuality, Honesty, and Consistency. Each section covers the corresponding research questions, related approaches, and evaluation settings, with links across dimensions where their concerns overlap.
Focus: the factual correctness of generated content, especially in knowledge-intensive responses.
Improving factuality involves understanding knowledge boundaries, detecting factual errors, improving responses through alignment and inference, and evaluating the resulting generations. The following reading lists cover these research questions, methods, and evaluation resources.
Studies of what a model knows, how reliably it can access that knowledge, and how prompting, retrieval, and fine-tuning affect its knowledge limits.
Methods for identifying factual errors or estimating hallucination risk using sampled responses, semantic uncertainty, internal states, and reasoning behavior. These approaches connect hallucination detection with probing a model's knowledge and uncertainty.
Training approaches that improve factual responses through data construction, supervised fine-tuning, factuality preferences, reinforcement learning, or knowledge distillation. This category also includes learning to refuse questions beyond the model's knowledge.
Approaches that improve factual generation through retrieval, prompting, sampling, decoding, activation intervention, or verification and correction. Retrieval systems that also require training are included here for their generation-time use of external evidence.
Benchmarks and evaluation methods for factual correctness, hallucination recognition, and factual support in generated text. The list includes question-answering datasets such as TriviaQA, SciQ, and Natural Questions, together with the work introducing the open-domain NQ setup. Coverage extends to short-form answers, long-form claims, and retrieval-grounded generation.
Focus: recognizing knowledge and capability limits, and faithfully communicating uncertainty about generated responses.
Honesty involves recognizing knowledge and capability limits, estimating and calibrating confidence, and appropriately expressing uncertainty. It is model-specific: evaluating honesty requires assessing how a model's expressed certainty and answering behavior relate to its own knowledge and capabilities.
Studies of a model's ability to assess what it knows, recognize insufficient information, false premises, or ill-posed problems, and decide when to seek external evidence. This includes identifying mathematical unsolvability and distinguishing it from the model's own capability limits. Related studies of known and unknown questions are also collected under Knowledge Boundary.
Methods for estimating uncertainty from token probabilities, verbalized confidence, sampled responses, semantic variation, and internal representations. The list also includes uncertainty signals used to guide learning and agent decisions. Applications to factual-error detection are collected under Hallucination Detection.
Overview of four families of confidence and uncertainty estimation methods—likelihood-based, prompting-based, sampling-based, and training-based—and their limitations.
Methods for aligning confidence estimates with observed correctness, including prompting, post-hoc recalibration, supervised learning, and reinforcement learning. Coverage extends from calibration foundations to long-form generation, confidence in reasoning steps and final answers, multi-turn conversations, and tool-using agents.
Research on communicating uncertainty in words or numbers, explaining uncertainty, asking for clarification, and appropriately declining to answer. The goal is to acknowledge limitations while retaining useful answers to answerable questions. Related refusal-training approaches, including R-Tuning and RL from knowledge feedback, appear under Factuality Alignment.
Surveys, benchmarks, and evaluation studies for self-knowledge, confidence estimation, calibration, and faithful uncertainty expression. Evaluation considers discrimination between correct and incorrect answers, calibration, appropriate abstention, and robustness across tasks and interaction settings. Evaluation settings include mathematical solvability, theorem-proving sycophancy, and multilingual confidence estimation.
Focus: agreement with relevant evidence, task constraints, and interaction history, including multi-turn dialogue, RAG, and search agents.
Maintaining consistency requires understanding contextual faithfulness and the effects of context interference, developing mitigation methods, and evaluating behavior across interactions. When sources conflict or new evidence becomes available, it also involves resolving conflicts and appropriately updating earlier claims.
Studies of how generated content relates to source material, task instructions, and interaction history. Coverage includes context–memory conflicts, conflicts among external sources, irrelevant-context interference, long-context information loss, and contradictions across conversation turns. These studies connect source faithfulness in summarization and dialogue to reliability in multi-turn agents.
Approaches for improving source faithfulness and coherence across interactions. Coverage includes key-information extraction and reranking, context compression, and prompting, together with training for context adherence, conflict-aware decoding, verification, and response correction. IRCoT, ReAct, Search-o1, Search-R1, and RAG-Gym extend this coverage to iterative retrieval, reasoning, and agent training. Related work on context-aware decoding, Self-RAG, and general retrieval-augmented generation is listed under Factuality Inference.
Benchmarks, metrics, and human-evaluation protocols for source support, contradiction detection, knowledge-conflict handling, conversational memory, and consistency across turns. Coverage spans summarization, data-to-text generation, knowledge-grounded dialogue, RAG, and agent interactions. Question-answering datasets provide task-level evaluation resources, complemented by checks of source support and consistency across turns. Evaluation considers evidence use and context adherence alongside answer correctness, task success, and interaction cost. Related resources, including Natural Questions, TriviaQA, RAGTruth, DeepTRACE, and long-document factual-consistency stress tests, are listed under Factuality Evaluation.
Contributions of relevant papers, surveys, benchmarks, and tools are welcome. Place each work under the subsection matching its main research question and cross-reference other dimensions when useful.
Use a linked title followed by the publication venue and year, matching the existing reading lists:
- [Paper title](PAPER_URL) — **Venue YYYY**.
Verify bibliographic details before adding an entry. Briefly describe its relevance when helpful, and include code or data links when available.
JavaScript
96.9%
CSS
2.2%
Factuality · Honesty · Consistency
A curated literature collection on building reliable large language models to mitigate hallucination.
Background · Factuality · Honesty · Consistency · Contribute
This project follows the research themes of my PhD thesis, A Framework of Building Reliable Large Language Models to Mitigate Hallucination, and curates related literature on factuality, honesty, and consistency.
Mitigating hallucination requires understanding both why models produce unreliable responses and how to improve their behavior. Reliability encompasses factual accuracy, awareness of knowledge limitations, and coherence with relevant context.
A reliable LLM should use knowledge to produce factually accurate responses, honestly communicate uncertainty and knowledge limitations, and maintain contextual consistency throughout multi-turn interactions.
| Dimension | Central question | Literature focus |
|---|---|---|
| Factuality | Does the response agree with verifiable facts? | Factual knowledge, knowledge boundaries, hallucination detection, and factuality improvement. |
| Honesty | Does the model accurately communicate what it knows and does not know? | Self-awareness, confidence estimation, calibration, and uncertainty expression. |
| Consistency | Does the response remain coherent with relevant evidence and the interaction history? | Contextual faithfulness, multi-turn dialogue, retrieval, and search agents. |
Research on hallucination, knowledge, and uncertainty is organized within these three dimensions. Papers are grouped by their primary research question, with cross-references for work spanning multiple dimensions.
Papers are grouped by their primary research focus and ordered by year, newest first. Each entry links to the paper and gives its conference, workshop, or journal and year. Papers with a verified acceptance but no linked proceedings version are marked accepted. Where a formal venue has not been verified, the entry is labeled arXiv preprint with its initial submission month. Bibliographic information was checked on 2026-09-22.
Large language models have advanced from text generation to agents that retrieve information, use tools, and interact with external environments. Yet they can still produce factually incorrect, fabricated, deceptive, or inconsistent content. These hallucination-related failures pose a central challenge to LLM reliability.
Such failures can mislead users even when the responses appear fluent and confident. The examples below illustrate this concern in health care and economics, motivating the need for reliable generation in knowledge-intensive applications.
Illustrative hallucinated responses in health care and economics; highlighted claims demonstrate unreliable outputs.
A central challenge is the gap between learning frequent textual patterns and reliably applying factual knowledge. During pre-training, next-token prediction teaches models to produce plausible continuations from large text corpora. During post-training, models learn response patterns from instruction and feedback data. Learning these patterns does not by itself ensure factual understanding or a clear boundary between known and unknown information.
The example below illustrates how a familiar statement about the first person on the Moon can be reused to answer a superficially similar question about Mars. This pattern-matching perspective connects the training process to the unreliable behaviors discussed next.
An illustration of hallucination arising from reliance on familiar textual patterns.
Unreliable behavior manifests at three levels: the model's expression of its own knowledge, the factual content of a single response, and the coherence of responses across an interaction.
Three manifestations of unreliability: dishonesty, non-factual responses, and contextual inconsistency.
These failure modes motivate complementary reliability requirements. Assessing answer correctness alone does not capture whether a model communicates its limitations appropriately or remains consistent throughout a conversation.
Research on hallucination mitigation spans data enhancement, fine-tuning, and retrieval-augmented generation (RAG). The causes and manifestations of hallucination motivate three complementary objectives for reliable LLMs:
A conceptual framework connecting hallucination causes, unreliable outputs, and the three dimensions of LLM reliability.
Following this progression, the literature collection is organized into Factuality, Honesty, and Consistency. Each section covers the corresponding research questions, related approaches, and evaluation settings, with links across dimensions where their concerns overlap.
Focus: the factual correctness of generated content, especially in knowledge-intensive responses.
Improving factuality involves understanding knowledge boundaries, detecting factual errors, improving responses through alignment and inference, and evaluating the resulting generations. The following reading lists cover these research questions, methods, and evaluation resources.
Studies of what a model knows, how reliably it can access that knowledge, and how prompting, retrieval, and fine-tuning affect its knowledge limits.
Methods for identifying factual errors or estimating hallucination risk using sampled responses, semantic uncertainty, internal states, and reasoning behavior. These approaches connect hallucination detection with probing a model's knowledge and uncertainty.
Training approaches that improve factual responses through data construction, supervised fine-tuning, factuality preferences, reinforcement learning, or knowledge distillation. This category also includes learning to refuse questions beyond the model's knowledge.
Approaches that improve factual generation through retrieval, prompting, sampling, decoding, activation intervention, or verification and correction. Retrieval systems that also require training are included here for their generation-time use of external evidence.
Benchmarks and evaluation methods for factual correctness, hallucination recognition, and factual support in generated text. The list includes question-answering datasets such as TriviaQA, SciQ, and Natural Questions, together with the work introducing the open-domain NQ setup. Coverage extends to short-form answers, long-form claims, and retrieval-grounded generation.
Focus: recognizing knowledge and capability limits, and faithfully communicating uncertainty about generated responses.
Honesty involves recognizing knowledge and capability limits, estimating and calibrating confidence, and appropriately expressing uncertainty. It is model-specific: evaluating honesty requires assessing how a model's expressed certainty and answering behavior relate to its own knowledge and capabilities.
Studies of a model's ability to assess what it knows, recognize insufficient information, false premises, or ill-posed problems, and decide when to seek external evidence. This includes identifying mathematical unsolvability and distinguishing it from the model's own capability limits. Related studies of known and unknown questions are also collected under Knowledge Boundary.
Methods for estimating uncertainty from token probabilities, verbalized confidence, sampled responses, semantic variation, and internal representations. The list also includes uncertainty signals used to guide learning and agent decisions. Applications to factual-error detection are collected under Hallucination Detection.
Overview of four families of confidence and uncertainty estimation methods—likelihood-based, prompting-based, sampling-based, and training-based—and their limitations.
Methods for aligning confidence estimates with observed correctness, including prompting, post-hoc recalibration, supervised learning, and reinforcement learning. Coverage extends from calibration foundations to long-form generation, confidence in reasoning steps and final answers, multi-turn conversations, and tool-using agents.
Research on communicating uncertainty in words or numbers, explaining uncertainty, asking for clarification, and appropriately declining to answer. The goal is to acknowledge limitations while retaining useful answers to answerable questions. Related refusal-training approaches, including R-Tuning and RL from knowledge feedback, appear under Factuality Alignment.
Surveys, benchmarks, and evaluation studies for self-knowledge, confidence estimation, calibration, and faithful uncertainty expression. Evaluation considers discrimination between correct and incorrect answers, calibration, appropriate abstention, and robustness across tasks and interaction settings. Evaluation settings include mathematical solvability, theorem-proving sycophancy, and multilingual confidence estimation.
Focus: agreement with relevant evidence, task constraints, and interaction history, including multi-turn dialogue, RAG, and search agents.
Maintaining consistency requires understanding contextual faithfulness and the effects of context interference, developing mitigation methods, and evaluating behavior across interactions. When sources conflict or new evidence becomes available, it also involves resolving conflicts and appropriately updating earlier claims.
Studies of how generated content relates to source material, task instructions, and interaction history. Coverage includes context–memory conflicts, conflicts among external sources, irrelevant-context interference, long-context information loss, and contradictions across conversation turns. These studies connect source faithfulness in summarization and dialogue to reliability in multi-turn agents.
Approaches for improving source faithfulness and coherence across interactions. Coverage includes key-information extraction and reranking, context compression, and prompting, together with training for context adherence, conflict-aware decoding, verification, and response correction. IRCoT, ReAct, Search-o1, Search-R1, and RAG-Gym extend this coverage to iterative retrieval, reasoning, and agent training. Related work on context-aware decoding, Self-RAG, and general retrieval-augmented generation is listed under Factuality Inference.
Benchmarks, metrics, and human-evaluation protocols for source support, contradiction detection, knowledge-conflict handling, conversational memory, and consistency across turns. Coverage spans summarization, data-to-text generation, knowledge-grounded dialogue, RAG, and agent interactions. Question-answering datasets provide task-level evaluation resources, complemented by checks of source support and consistency across turns. Evaluation considers evidence use and context adherence alongside answer correctness, task success, and interaction cost. Related resources, including Natural Questions, TriviaQA, RAGTruth, DeepTRACE, and long-document factual-consistency stress tests, are listed under Factuality Evaluation.
Contributions of relevant papers, surveys, benchmarks, and tools are welcome. Place each work under the subsection matching its main research question and cross-reference other dimensions when useful.
Use a linked title followed by the publication venue and year, matching the existing reading lists:
- [Paper title](PAPER_URL) — **Venue YYYY**.
Verify bibliographic details before adding an entry. Briefly describe its relevance when helpful, and include code or data links when available.
JavaScript
96.9%
CSS
2.2%