ybwang119/Awesome-reasoning-safety

This repo is for the safety topic, including attacks, defenses and studies related to reasoning and RL

67

17 commits

updated Sep 5, 2025

See the code

README

Awesome-reasoning-safety

Awesome Maintenance Last Commit CC BY 4.0

This repo is for the trustworthy topics in reasoning-related technique, including but not limited to attacks, defenses, studies and benchmarks related to CoT, reasoning and RL. Data are mainly from arxiv.

๐Ÿš€ Our Survey

๐Ÿ“ข Check out our survey paper: A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

๐Ÿ“– Citation

If you find our survey useful, please cite it as:

@article{wang2025comprehensive,
  title   = {A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models},
  author  = {Wang, Yanbo and Yu, Yongcan and Liang, Jian and He, Ran},
  journal = {arXiv preprint arXiv:2509.03871},
  year    = {2025}
}

๐Ÿค Contribute

We welcome contributions from the community, including adding your papers, modifying topics, or other comments related to our survey or repo!๐ŸŽ‰

  • Found a missing paper?
  • Have suggestions to improve this list?

๐Ÿ‘‰ Feel free to open an issue or submit a pull request.

If you like this project, donโ€™t forget to โญ๏ธ star it โ€” it helps more people discover it! ๐Ÿ’–


๐Ÿ“šTable of contents


Truthfulness: Hallucination, reasoning faithfulness

Hallucination

Click to hide/show the paper list
TitleVenueDatetopicCode
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLMarxiv08/27, 2025hallucination study of LRM-
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMsarxiv08/26, 2025evaluation of quantized LLMGithub
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QAarxiv08/19, 2025hallucination mitigation with reasoning-
Mitigating Hallucinations in Large Language Models via Causal Reasoningarxiv08/17, 2025hallucination mitigation with reasoningGithub
Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysisarxiv08/06, 2025hallucination study of LRM-
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMsarxiv08/06, 2025hallucination mitigation with reasoning-
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processesarxiv08/02, 2025hallucination mitigation with reasoning-
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraintsarxiv07/22, 2025calibration improvement with reasoning-
Beyond Binary Rewards: Training LMs to Reason About Their Uncertaintyarxiv07/22, 2025calibration improvement with reasoning-
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factualityarxiv06/24, 2025hallucination mitigation with reasoningGithub
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Modelsarxiv06/23, 2025evaluation on math of LRMGithub
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?arxiv06/22, 2025evaluation on uncertainty of LRM-
Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluationarxiv06/20, 2025evaluation on hallucination detectionGithub
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detectionarxiv06/05, 2025evaluation on detection with reasoning-
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Modelsarxiv06/05, 2025evaluation on hallucination detection of LRMGithub
The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Modelsarxiv05/30, 2025evaluation of LRM-
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMarxiv05/30, 2025evaluation and mitigation with reaosning-
Are Reasoning Models More Prone to Hallucination?arxiv05/29, 2025evaluation of LRM-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Modelsarxiv05/23, 2025evaluation of MLRMProject
Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical ReasoningJSAI-isAI 202505/23, 2025evaluation on math with reasoningProject
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluationarxiv05/20, 2025unanswerable datasetGitHub
The Hallucination Tax of Reinforcement Finetuningarxiv05/20, 2025hallucination study with RLHuggingface
Toward Reliable Biomedical Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Modelsarxiv05/20, 2025hallucination detection with reasoningGitHub
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspectivearxiv05/19, 2025hallucination study of LRMAnonymous
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Modelsarxiv05/19, 2025hallucination study of LRMAnonymous
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Featuresarxiv05/17, 2025hallucination study of LRM-
Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detectionarxiv04/13, 2025hallucination detection with reasoning-
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoningarxiv04/08, 2025hallucinationn detection with reasoning-
Do Chains-of-Thoughts of Large Language Models Suffer from Hallucinations, Cognitive Biases, or Phobias in Bayesian Reasoning?arxiv03/19, 2025hallucination study with reasoning-
Grounded Chain-of-Thought for Multimodal Large Language Modelsarxiv03/17, 2025hallucination study with reasoningGithub
Mitigating reasoning hallucination through Multi-agent Collaborative FilteringESWA 202503/05, 2025hallucination mitigation with reasoning-
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinkingarxiv01/02, 2025hallucination mitigation with reasoning-
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucinationarxiv11/15, 2024hallucination mitigation with reasoningGithub
HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought ReasoningEMNLP 202411/12, 2024hallucination measurement with reasoning-
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoningarxiv10/08, 2024hallucination detection with reasoningAnonymous
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generationarxiv06/17, 2024hallucination mitigation with reasoningGithub

Reasoning Faithfulness

Click to hide/show the paper list
TitleVenueDatetopicCode
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in ExplanationsICML 2025 Workshop on Actionable Interpretability07/15, 2025faithfulness improvementGithub
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoningarxiv07/13, 2025faithfulness improvementProject
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoningarxiv06/28, 2025faithfulness improvement-
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoningarxiv06/24, 2025faithfulness understanding-
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?arxiv06/13, 2025MLLM faithfulness evaluationGithub
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learningarxiv05/22, 2025faitufulness improvement-
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Modelsarxiv05/20, 2025faithfulness measurementGithub
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanationsarxiv05/20, 2025faithfulness measurement-
Reasoning Models Donโ€™t Always Say What They Thinkwebsite04/03, 2025faithfulness understanding-
Reasoning Inconsistencies and How to Mitigate Them in Deep Learningarxiv04/03, 2025faithfulness improvement-
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Modelsarxiv03/28, 2025faithfulness understandingGithub
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Modelsarxiv03/14, 2025faithfulness improvement-
Chain-of-Thought Reasoning In The Wild Is Not Always Faithfularxiv03/13, 2025faithfulness understanding-
A Causal Lens for Evaluating Faithfulness Metricsarxiv02/26, 2025faithfulness measurement-
Measuring faithfulness of chains of thought by unlearning reasoning stepsarxiv02/20, 2025faithfulness measurementGithub
Are DeepSeek R1 And Other Reasoning Models More Faithful?ICLR 2025 Workshop01/14, 2025faithfulness understanding-
Graph-Guided Textual Explanation Generation Frameworkarxiv12/16, 2024faithfulness improvement-
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processingarxiv11/27, 2024faithfulness understanding-
On the Impact of Fine-Tuning on Chain-of-Thought Reasoningarxiv11/22, 2024faithfulness understanding-
Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question AnsweringACM MM 2410/28, 2024faithfulness improvement-
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Modelsarxiv10/18, 2024faithfulness improvement-
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External ContextsICLR 202410/18, 2024faithfulness improvementGithub
Enhancing Large Language Models' Situated Faithfulness to External Contextsarxiv10/18, 2024faithfulness improvement-
FLARE: Faithful Logic-Aided Reasoning and Explorationarxiv10/14, 2024faithfulness improvement-
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoningarxiv10/14, 2024faithfulness improvementGithub
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language ModelsICML 2024 Workshop06/15, 2024faithfulness understanding-
XPrompt:Explaining Large Language Model's Generation via Joint Prompt Attributionarxiv05/30, 2024faithfulness understanding-
Faithful Logical Reasoning via Symbolic Chain-of-ThoughtACL 202405/28, 2024faithfulness improvement-
Dissociation of Faithful and Unfaithful Reasoning in LLMsarxiv05/23, 2024faithfulness understandingGithub
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question AnsweringACL 202505/22, 2024faithfulness improvenmentGithub
Faithful Reasoning over Scientific ClaimsAAAI 202405/20, 2024faithfulness improvement-
Towards Better Chain-of-Thought: A Reflection on Effectiveness and FaithfulnessACL 2025 Findings05/29, 2024faithfulness improvementGithub
Markovian Transformers for Informative Language Modelingarxiv04/29, 2024faithfulness improvement-
Fact: Teaching MLLMs with Faithful, Concise and Transferable RationalesACM MM 202404/17, 2024faithfulness improvement-
Argumentative Large Language Models for Explainable and Contestable Claim VerificationAAAI 202504/11, 2024faithfulness improvementGithub
Recent Developments on Accountability and Explainability for Complex Reasoning TasksSpringer04/06, 2024faithfulness improvement-
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language ModelsACL 202404/04, 2024faithfulness measurement-
How Likely Do LLMs with CoT Mimic Human Reasoning?arxiv02/25, 2024faithfulness understandingGithub
Chain-of-Thought Unfaithfulness as Disguised AccuracyTMLR02/22, 2024faithfulness understanding-
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought ReasoningEMNLP 2024 Findings02/21, 2024faithfulness improvementGithub
How Interpretable are Reasoning Explanations from Prompting Large Language Models?NAACL 2024 Findings02/19, 2024faithfulness understandingGithub
FaithLM: Towards Faithful Explanations for Large Language ModelsEMNLP 2024 Findings02/07, 2024faithfulness improvement-
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Modelsarxiv02/07, 2024faithfulness understanding-
Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningICLR 202410/02, 2023faithfulness improvementGithub
Measuring Faithfulness in Chain-of-Thought Reasoningarxiv07/17, 2023faithfulness measurement-
Question Decomposition Improves the Faithfulness of Model-Generated Reasoningarxiv07/17, 2023faithfulness improvementGithub
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical ReasoningEMNLP 2023 Findings05/20, 2023faithfulness improvementGithub
Language Models Donโ€™t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Promptingarxiv05/07, 2023faithfulness understanding-
Faithful Chain-of-Thought ReasoningIJCNLP-AACL 202301/31, 2023faithfulness improvementGithub

Safety: Assessment, Jailbreak, Alignment, Backdoor

Vulnerability assessment

Click to hide/show the paper list
TitleVenueDatetopicCode
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusalarxiv08/15, 2025over-refusal evaluationGithub
Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language ModelsMM 202508/12, 2025datasetGithub
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Modelsarxiv08/06, 2025content-safety evaluationGithub
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability ComponentsICML 2025 TAIG workshop07/10, 2025agenic safety evaluation-
FORTRESS: Frontier Risk Evaluation for National Security and Public Safetyarxiv06/24, 2025safety evaluationHuggingface
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning ModelsLLMSEC 202506/16, 2025content-safety evaluation-
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own ReasoningICML 202506/06, 2025agenic safety evaluationGithub
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Modelsarxiv05/26, 2025content-safety evaluationGithub
RRTL: Red Teaming Reasoning Large Language Models in Tool Learningarxiv05/21, 2025tool-learning safety evaluation-
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoningarxiv04/02, 2025content-safety evaluation-
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findingsarxiv03/19, 2025content-safety evaluation-
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contextsarxiv03/18, 2025content-safety evaluation-
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectivesarxiv03/13, 2025multi-lingual safety evaluation-
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1arxiv02/27, 2025content-safety evaluation-
Evaluating security risk in deepseek and other frontier reasoning modelswebsite02/26, 2025content-safety evaluation-
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuningarxiv02/21, 2025content-safety evaluation-
o3-mini vs DeepSeek-R1: Which One is Safer?arxiv01/30, 2025content-safety evaluation-
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategiesarxiv01/28, 2025content-safety evaluation-

Jailbreak

Jailbreak attack

Click to hide/show the paper list
TitleVenueDatetopicCode
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Promptsarxiv08/14, 2025jailbreak with CoTGithub
The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?arxiv08/05, 2025multi-modal jailbreak attack-
Adversarial Manipulation of Reasoning Models using Internal RepresentationsICML 2025 R2-FM Workshop07/01, 2025White-box attackGithub
HauntAttack: When Attack Follows Reasoning as a Shadowarxiv06/08, 2025jailbreak attack-
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Modelsarxiv05/26, 2025jailbreak attack-
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Modelsarxiv05/26, 2025jailbreak MLRM-
Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?arxiv05/23, 2025jailbreak analysis-
Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Modelsarxiv05/23, 2025jailbreak with CoT-
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphersarxiv05/22, 2025jailbreak attack-
AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Modelsarxiv05/16, 2025jailbreak attackGithub
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challengesarxiv04/27, 2025jailbreak attackProject
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphersarxiv03/16, 2025jailbreak with CoTGithub
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Modelsarxiv03/11, 2025jailbreak with CoT-
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaosarxiv02/19, 2025jailbreak attack-
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinkingarxiv02/18, 2025jailbreak attack-
Adversarial Reasoning at Jailbreaking Timearxiv02/03, 2025jailbreak with CoT-
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled modelsarxiv02/03, 2025finetuning attack-

Jailbreak defense

Click to hide/show the paper list
TitleVenueDatetopicCode
IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinementarxiv08/27, 2025guardrail model-
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Languagearxiv08/21, 2025toxicity detection with CoT-
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questionsarxiv08/08, 2025Jailbreak defenseGithub
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Momentsarxiv08/06, 2025guardrail model-
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulatorarxiv07/18, 2025agenic safetyproject
Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Modelsarxiv07/16, 2025CoT Monitor-
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv07/15, 2025CoT Monitor-
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv07/07, 2025CoT Monitor-
AgentBreeder: Mitigating the AI Safety Impact of Multi-Agent Scaffolds via Self-ImprovementICLR 202506/25, 2025Agentic safetyGithub
Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoningarxiv06/10, 2025harmful meme detectionLink
RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsarxiv06/09, 2025guardrail modelGithub
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Modelsarxiv05/26, 2025guardrail model-
ReasoningShield: Content Safety Detection over Reasoning Traces of Large Reasoning Modelsarxiv05/22, 2025Guardrail modelGithub
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimizationarxiv05/20, 2025LLM-as-a-judgeGithub
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMsarxiv05/20, 2025guardrail model-
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoningarxiv05/16, 2025guardrail model-
On the Robustness of Reward Models for Language Model Alignmentarxiv05/12, 2025reward model-
LlamaFirewall: An open source guardrail system for building secure AI agentsarxiv05/06, 2025Guardrail model-
MR. Guard: Multilingual Reasoning Guardrail using Curriculum Learningarxiv04/21, 2025Guardrail model-
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimizationarxiv04/17, 2025multi-modal guardrail model-
X-Guard: Multilingual guard agent for content moderationarxiv04/11, 2025multi-lingual guardrail modelGithub
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Controlarxiv04/23, 2025jailbreak defense studyGithub
ShieldAgent: Shielding Agents via Verifiable Safety Policy ReasoningICML 202503/26, 2025guardrail modelGithub
GuardReasoner: Towards Reasoning-based LLM Safeguardsarxiv01/30, 2025guardrail modelGithub
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled ReasoningICML 202506/13, 2024guardrail modelProject
$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical ReasoningICLR 2025 spotlight06/08, 2024Guardrail modelGithub

Alignment

Click to hide/show the paper list
TitleVenueDatetopicCode
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodalityarxiv08/25, 2025VLM alignment with reasoningGithub
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigationarxiv08/21, 2025alignment with reasoning-
SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language ModelsEMNLP 202508/21, 2025alignment with reasoningGithub
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Alignerarxiv08/20, 2025Inference-time alignment with TTS-
FuSaR: A Fuzzification-Based Method for LRM Safety-Reasoning Balancearxiv08/18, 2025alignment of LRMGithub
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Modelsarxiv08/08, 2025alignment with reasoningGithub
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledgearxiv08/01, 2025alignment of LRMGithub
UnsafeChain: Enhancing Reasoning Model Safety via Hard Casesarxiv07/29, 2025alignment of LRMGithub
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Lawarxiv07/24, 2025alignment of LRM-
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learningarxiv07/20, 2025alignment with reasoningGithub
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoningarxiv07/14, 2025alignment with reasoningProject
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approacharxiv07/03, 2025test-time alignment with RLGithub
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulnessarxiv07/02, 2025VLM alignment with reasoningGithub
Reasoning as an Adaptive Defense for Safetyarxiv07/01, 2025alignment with reasoningProject
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Modelsarxiv06/23, 2025multimodal alignment datasetHuggingface
How Alignment Shrinks the Generative Horizonarxiv06/21, 2025alignment study of LRMProject
SafeCoT: Improving VLM Safety with Minimal Reasoningarxiv06/09, 2025VLM alignment with reasoning-
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Modelsarxiv06/09, 2025alignment with reasoningGithub
Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurancearxiv06/06, 2025alignment with reasoningGithub
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignmentarxiv06/01, 2025alignment with reasoning-
Mitigating Deceptive Alignment via Self-Monitoringarxiv05/24, 2025alignment of LRMProject
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoningarxiv05/22, 2025alignment of LRMProject
From Evaluation to Defense: Advancing Safety in Video Large Language Modelsarxiv05/22, 2025multimodal alignment with reasoning-
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Studyarxiv05/21, 2025alignment study of LRMGitHub
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learningarxiv05/20, 2025alignment of LRM-
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignmentarxiv05/20, 2025alignment of LRMProject
Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correctionarxiv05/19, 2025agent alignment with reasoningHuggingface
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured ReasoningCoLM 202505/12, 2025over reject mitigation with reasoningProject
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Modelarxiv05/10, 2025alignment of MLRMGithub
Energy-Based Reward Models for Robust Language Model AlignmentCOLM 202504/17, 2025reward model improvementGithub
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv04/17, 2025CoT monitor in alignment of LRM-
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capabilityarxiv04/14, 2025alignment of LRMHuggingface
SaRO: Enhancing LLM Safety through Reasoning-based Alignmentarxiv04/13, 2025alignment with reasoningGithub
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Modelsarxiv04/09, 2025alignment of MLRMGithub
ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimizationarxiv04/06, 2025alignment with reasoning-
STAR-1: Safer Alignment of Reasoning LLMs with 1K Dataarxiv04/02, 2025alignment of LRMProject
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignmentarxiv03/23, 2025reward model improvement-
Think Before Refusal: Triggering Safety Reflection in LLMs to Mitigate False Refusal Behaviorarxiv03/22, 2025over-reject mitigation with reasoning-
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safetyarxiv03/06, 2025alignment with reasoning-
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonablearxiv03/01, 2025alignment study of LRMGithub
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreakingarxiv02/18, 2025alignment with reasoningGithub
Safechain: Safety of language models with long chain-of-thought reasoning capabilitiesarxiv02/17, 2025alignment of LRMProject
Safety Reasoning with GuidelinesICML 202502/06, 2025alignment with reasoning-
STAIR: Improving Safety Alignment with Introspective Reasoningarxiv02/04, 2025alignment with reasoningGithub
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peersarxiv02/01, 2025alignment study of LRM-
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoningarxiv01/31, 2025alignment with reasoninganonymous
Backtracking Improves Generation SafetyICLR 2025 Oral09/22, 2024alignment with reasoning-

Backdoor

Click to hide/show the paper list
TitleVenueDatetopicCode
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIsarxiv08/08, 2025injection defenseGithub
Defend LLMs Through Self-ConsciousnessKDD 2025 Workshop08/04, 2025injection defense-
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profitarxiv07/24, 2025training-phase data poisoningGithub
Thought Purity: Defense Paradigm For Chain-of-Thought Attackarxiv07/16, 2025injection defense-
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Modelsarxiv06/16, 2025Training-time
data poisoningHuggingface
GUARD:Dual-Agent based Backdoor Defense on Chain-of-Thought in Neural Code Generationarxiv05/27, 2025backdoor defenseGithub
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systemsarxiv05/22, 2025inference-time prompt manipulation-
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injectionarxiv05/10, 2025inference-time prompt manipulation-
Practical Reasoning Interruption Attacks on Reasoning Large Language Modelsarxiv05/10, 2025inference-time prompt manipulation-
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compressionarxiv04/29, 2025inference-time prompt manipulation-
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMsarxiv04/08, 2025training-phase data poisoning-
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detectionarxiv03/27, 2025injection defense-
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Modelsarxiv02/16, 2025training-phase data poisoningGithub
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMsarxiv01/24, 2025inference-time prompt manipulation-
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language ModelsACL 2025 Findings12/20, 2024backdoor defense-
SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generationarxiv12/08, 2025training-phase data poisoning-
BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsICLR 202401/20, 2024inference-time prompt manipulation-

Robustness: evaluation and attack, improvement, overthinking, underthinking

Evaluation and attack

Click to hide/show the paper list
TitleVenueDatetopicCode
Measuring Sycophancy of Language Models in Multi-turn Dialoguesarxiv08/25, 2025sycophancy evaluationGithub
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Frameworkarxiv08/05, 2025evaluation on code-
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questionsarxiv07/29, 2025evaluation on unanswerable questionGithub
When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMsarxiv07/22, 2025adversarial attack-
Reasoning Models Are More Easily Gaslighted Than You Thinkarxiv06/11, 2025evaluation on misleading inputs-
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questionsarxiv06/10,2025evaluation on unanswerable questionGithub
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generationarxiv06/07, 2025evaluation on adversarial promptingGithub
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsarxiv05/30, 2025Failure analysis-
CodeCrash: Stress Testing LLM Reasoning under Structural and Semantic Perturbationsarxiv05/23, 2025evalution on code comprehensionGitHub
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contextsarxiv04/28, 2025evaluation on math[Github](https://github .com/QwenLM/PolyMath)
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Stepsarxiv03/25, 2025adversarial attack-
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1arxiv03/13, 2025adversarial attackGithub
Benchmarking Reasoning Robustness in Large Language Modelsarxiv03/06, 2025evaluation-
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning ModelsCOLM 202503/03, 2025adversarial attackHuggingface
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?NeurIPS 202402/27, 2025evaluation on self-reflectionGithub
A Closer Look at System Prompt Robustnessarxiv02/15, 2025evaluation on system promptGithub
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbationsarxiv02/10, 2025evaluation on math-
Understanding the Dark Side of LLMs' Intrinsic Self-Correctionarxiv12/19, 2024self reflection studyGithub
Stepwise Reasoning Disruption Attack of LLMsACL 202512/16, 2024adversarial attackGithub
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?arxiv10/31, 2024evaluation on noisy rationalesGithub
RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Modelsarxiv06/16, 2024evaluationGithub
Preemptive Answer "Attacks" on Chain-of-Thought ReasoningFindings of ACL 202405/31, 2024adversarial attack-

Improvement

Click to hide/show the paper list

Overthinking and underthinking

Click to hide/show the paper list
TitleVenueDatetopicCode
Don't Think Twice! Over-Reasoning Impairs Confidence CalibrationICML 2025 Workshop08/20, 2025overthinking study-
OptimalThinkingBench: Evaluating Over and Underthinking in LLMsarxiv08/18, 2025evaluationGithub
Excessive Reasoning Attack on Reasoning LLMsarxiv06/17, 2025overthinking attack-
Mitigating Overthinking in Large Reasoning Models via Manifold Steeringarxiv05/28, 2025overthinking mitigation-
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Modelsarxiv05/27, 2025underthinking mitigationGithub
Internal Bias in Reasoning Models leads to Overthinkingarxiv05/22, 2025overthinking study-
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMsarxiv04/30, 2025reasoning length study-
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Criticalarxiv04/11, 2025overthinking study-
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilitiesarxiv03/23, 2025reasoning length studyGuthub
DNA Bench: When Silence is Smarter --Benchmarking Over-Reasoning in Reasoning LLMsarxiv03/19, 2025overthinking evaluation-
Output Length Effect on DeepSeek-R1's Safety in Forced Thinkingarxiv03/02, 2025reasoning length study-
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasksarxiv02/12, 2025overthinking studyGithub
OverThink: Slowdown Attacks on Reasoning LLMsarxiv02/04, 2025overthinking attackGithub
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMsarxiv01/30, 2025underthinking study-
Large Language Models Struggle with Unreasonability in Math Problemsarxiv03/28, 2024overthinking study-

Privacy

Click to hide/show the paper list

Fairness

Click to hide/show the paper list
TitleVenueDatetopicCode
Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languagesarxiv08/22, 2025multi-lingual bias-
Do Biased Models Have Biased Thoughts?COLM 202508/08, 2025study-
FairReason: Balancing Reasoning and Social Bias in MLLMsarxiv07/30, 2025fairness improvementGithub
The Emperor's New Chain-of-Thought: Probing Reasoning Theater Bias in Large Reasoning Modelsarxiv07/18, 2025fairness evaluation-
Guiding LLM Decision-Making with Fairness Reward Modelsarxiv07/15, 2025fairness improvementGithub
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Modelsarxiv07/03, 2025fairness evaluation-
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoningarxiv06/24, 2025persona biasGithub
Detection, Classification, and Mitigation of Gender Bias in Large Language Modelsarxiv06/14, 2025gender bias-
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasksarxiv06/09, 2025dialect fairnessGithub
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language ModelsACL 2025 Findings04/30, 2025bias detection-
Prompting techniques for reducing social bias in LLMs through system 1 and system 2 cognitive processesarxiv04/26, 2024social biasGithub
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsICLR 202411/08, 2023reasoning biasProject
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgptNature Computational Science10/05, 2023reasoning bias-
Click to hide/show the paper list
Click to hide/show the paper list
TitleVenueDatetopicCode
It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topicsarxiv08/18, 2025persuasion evaluation-
Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillationarxiv08/15, 2025safe data distillation-
Representation Bending for Large Language Model Safetyarxiv07/14, 2025safety-
Defending Against Prompt Injection With a Few Defensive Tokensarxiv07/10, 2025prompt injection defense-
Truth Neuronsarxiv07/08, 2025interpretability study-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMsarxiv06/06, 2025safety evaluation-
Give Me FP32 or Give Me Death? Challenges and Solutions for Reproducible Reasoningarxiv06/11, 2025reasoning reproducibility-
If Pigs Could Fly... Can LLMs Logically Reason Through Counterfactuals?arxiv05/28, 2025adversarial evaluationAnonymous
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoningarxiv05/28, 2025reasoning analysis-
Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphsarxiv05/21, 2025reasoning on knowledge graphs-
Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Modelsarxiv05/20, 2025alignment studyGitHub
Persuasion and Safety in the Era of Generative AIarxiv05/18, 2025persuasion study-
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM SecurityICLR 2025 Workshop04/29, 2025alignment study with reasoningGithub
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMsarxiv04/28, 2025moral evaluation-
Effectively Controlling Reasoning Models through Thinking Interventionarxiv03/31, 2025reasoning control-
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Modelsarxiv03/14, 2025reasoning control-
Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Modelsarxiv12/30, 2024hallucination study-
Alignment faking in large language modelsarxiv12/19, 2024alignment study of LLM-

Contributors

yuyongcan

10 commits

ybwang119

7 commits

ybwang119/Awesome-reasoning-safety

This repo is for the safety topic, including attacks, defenses and studies related to reasoning and RL

67

17 commits

updated Sep 5, 2025

See the code

README

Awesome-reasoning-safety

Awesome Maintenance Last Commit CC BY 4.0

This repo is for the trustworthy topics in reasoning-related technique, including but not limited to attacks, defenses, studies and benchmarks related to CoT, reasoning and RL. Data are mainly from arxiv.

๐Ÿš€ Our Survey

๐Ÿ“ข Check out our survey paper: A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

๐Ÿ“– Citation

If you find our survey useful, please cite it as:

@article{wang2025comprehensive,
  title   = {A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models},
  author  = {Wang, Yanbo and Yu, Yongcan and Liang, Jian and He, Ran},
  journal = {arXiv preprint arXiv:2509.03871},
  year    = {2025}
}

๐Ÿค Contribute

We welcome contributions from the community, including adding your papers, modifying topics, or other comments related to our survey or repo!๐ŸŽ‰

  • Found a missing paper?
  • Have suggestions to improve this list?

๐Ÿ‘‰ Feel free to open an issue or submit a pull request.

If you like this project, donโ€™t forget to โญ๏ธ star it โ€” it helps more people discover it! ๐Ÿ’–


๐Ÿ“šTable of contents


Truthfulness: Hallucination, reasoning faithfulness

Hallucination

Click to hide/show the paper list
TitleVenueDatetopicCode
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLMarxiv08/27, 2025hallucination study of LRM-
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMsarxiv08/26, 2025evaluation of quantized LLMGithub
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QAarxiv08/19, 2025hallucination mitigation with reasoning-
Mitigating Hallucinations in Large Language Models via Causal Reasoningarxiv08/17, 2025hallucination mitigation with reasoningGithub
Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysisarxiv08/06, 2025hallucination study of LRM-
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMsarxiv08/06, 2025hallucination mitigation with reasoning-
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processesarxiv08/02, 2025hallucination mitigation with reasoning-
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraintsarxiv07/22, 2025calibration improvement with reasoning-
Beyond Binary Rewards: Training LMs to Reason About Their Uncertaintyarxiv07/22, 2025calibration improvement with reasoning-
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factualityarxiv06/24, 2025hallucination mitigation with reasoningGithub
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Modelsarxiv06/23, 2025evaluation on math of LRMGithub
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?arxiv06/22, 2025evaluation on uncertainty of LRM-
Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluationarxiv06/20, 2025evaluation on hallucination detectionGithub
CLATTER: Comprehensive Entailment Reasoning for Hallucination Detectionarxiv06/05, 2025evaluation on detection with reasoning-
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Modelsarxiv06/05, 2025evaluation on hallucination detection of LRMGithub
The Hallucination Dilemma: Factuality-Aware Reinforcement Learning for Large Reasoning Modelsarxiv05/30, 2025evaluation of LRM-
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLMarxiv05/30, 2025evaluation and mitigation with reaosning-
Are Reasoning Models More Prone to Hallucination?arxiv05/29, 2025evaluation of LRM-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Modelsarxiv05/23, 2025evaluation of MLRMProject
Analyzing Logical Fallacies in Large Language Models: A Study on Hallucination in Mathematical ReasoningJSAI-isAI 202505/23, 2025evaluation on math with reasoningProject
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluationarxiv05/20, 2025unanswerable datasetGitHub
The Hallucination Tax of Reinforcement Finetuningarxiv05/20, 2025hallucination study with RLHuggingface
Toward Reliable Biomedical Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Modelsarxiv05/20, 2025hallucination detection with reasoningGitHub
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspectivearxiv05/19, 2025hallucination study of LRMAnonymous
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Modelsarxiv05/19, 2025hallucination study of LRMAnonymous
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Featuresarxiv05/17, 2025hallucination study of LRM-
Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detectionarxiv04/13, 2025hallucination detection with reasoning-
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoningarxiv04/08, 2025hallucinationn detection with reasoning-
Do Chains-of-Thoughts of Large Language Models Suffer from Hallucinations, Cognitive Biases, or Phobias in Bayesian Reasoning?arxiv03/19, 2025hallucination study with reasoning-
Grounded Chain-of-Thought for Multimodal Large Language Modelsarxiv03/17, 2025hallucination study with reasoningGithub
Mitigating reasoning hallucination through Multi-agent Collaborative FilteringESWA 202503/05, 2025hallucination mitigation with reasoning-
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinkingarxiv01/02, 2025hallucination mitigation with reasoning-
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucinationarxiv11/15, 2024hallucination mitigation with reasoningGithub
HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought ReasoningEMNLP 202411/12, 2024hallucination measurement with reasoning-
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoningarxiv10/08, 2024hallucination detection with reasoningAnonymous
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generationarxiv06/17, 2024hallucination mitigation with reasoningGithub

Reasoning Faithfulness

Click to hide/show the paper list
TitleVenueDatetopicCode
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in ExplanationsICML 2025 Workshop on Actionable Interpretability07/15, 2025faithfulness improvementGithub
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoningarxiv07/13, 2025faithfulness improvementProject
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoningarxiv06/28, 2025faithfulness improvement-
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoningarxiv06/24, 2025faithfulness understanding-
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?arxiv06/13, 2025MLLM faithfulness evaluationGithub
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learningarxiv05/22, 2025faitufulness improvement-
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Modelsarxiv05/20, 2025faithfulness measurementGithub
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanationsarxiv05/20, 2025faithfulness measurement-
Reasoning Models Donโ€™t Always Say What They Thinkwebsite04/03, 2025faithfulness understanding-
Reasoning Inconsistencies and How to Mitigate Them in Deep Learningarxiv04/03, 2025faithfulness improvement-
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Modelsarxiv03/28, 2025faithfulness understandingGithub
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Modelsarxiv03/14, 2025faithfulness improvement-
Chain-of-Thought Reasoning In The Wild Is Not Always Faithfularxiv03/13, 2025faithfulness understanding-
A Causal Lens for Evaluating Faithfulness Metricsarxiv02/26, 2025faithfulness measurement-
Measuring faithfulness of chains of thought by unlearning reasoning stepsarxiv02/20, 2025faithfulness measurementGithub
Are DeepSeek R1 And Other Reasoning Models More Faithful?ICLR 2025 Workshop01/14, 2025faithfulness understanding-
Graph-Guided Textual Explanation Generation Frameworkarxiv12/16, 2024faithfulness improvement-
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processingarxiv11/27, 2024faithfulness understanding-
On the Impact of Fine-Tuning on Chain-of-Thought Reasoningarxiv11/22, 2024faithfulness understanding-
Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question AnsweringACM MM 2410/28, 2024faithfulness improvement-
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Modelsarxiv10/18, 2024faithfulness improvement-
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External ContextsICLR 202410/18, 2024faithfulness improvementGithub
Enhancing Large Language Models' Situated Faithfulness to External Contextsarxiv10/18, 2024faithfulness improvement-
FLARE: Faithful Logic-Aided Reasoning and Explorationarxiv10/14, 2024faithfulness improvement-
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoningarxiv10/14, 2024faithfulness improvementGithub
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language ModelsICML 2024 Workshop06/15, 2024faithfulness understanding-
XPrompt:Explaining Large Language Model's Generation via Joint Prompt Attributionarxiv05/30, 2024faithfulness understanding-
Faithful Logical Reasoning via Symbolic Chain-of-ThoughtACL 202405/28, 2024faithfulness improvement-
Dissociation of Faithful and Unfaithful Reasoning in LLMsarxiv05/23, 2024faithfulness understandingGithub
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question AnsweringACL 202505/22, 2024faithfulness improvenmentGithub
Faithful Reasoning over Scientific ClaimsAAAI 202405/20, 2024faithfulness improvement-
Towards Better Chain-of-Thought: A Reflection on Effectiveness and FaithfulnessACL 2025 Findings05/29, 2024faithfulness improvementGithub
Markovian Transformers for Informative Language Modelingarxiv04/29, 2024faithfulness improvement-
Fact: Teaching MLLMs with Faithful, Concise and Transferable RationalesACM MM 202404/17, 2024faithfulness improvement-
Argumentative Large Language Models for Explainable and Contestable Claim VerificationAAAI 202504/11, 2024faithfulness improvementGithub
Recent Developments on Accountability and Explainability for Complex Reasoning TasksSpringer04/06, 2024faithfulness improvement-
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language ModelsACL 202404/04, 2024faithfulness measurement-
How Likely Do LLMs with CoT Mimic Human Reasoning?arxiv02/25, 2024faithfulness understandingGithub
Chain-of-Thought Unfaithfulness as Disguised AccuracyTMLR02/22, 2024faithfulness understanding-
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought ReasoningEMNLP 2024 Findings02/21, 2024faithfulness improvementGithub
How Interpretable are Reasoning Explanations from Prompting Large Language Models?NAACL 2024 Findings02/19, 2024faithfulness understandingGithub
FaithLM: Towards Faithful Explanations for Large Language ModelsEMNLP 2024 Findings02/07, 2024faithfulness improvement-
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Modelsarxiv02/07, 2024faithfulness understanding-
Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningICLR 202410/02, 2023faithfulness improvementGithub
Measuring Faithfulness in Chain-of-Thought Reasoningarxiv07/17, 2023faithfulness measurement-
Question Decomposition Improves the Faithfulness of Model-Generated Reasoningarxiv07/17, 2023faithfulness improvementGithub
Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical ReasoningEMNLP 2023 Findings05/20, 2023faithfulness improvementGithub
Language Models Donโ€™t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Promptingarxiv05/07, 2023faithfulness understanding-
Faithful Chain-of-Thought ReasoningIJCNLP-AACL 202301/31, 2023faithfulness improvementGithub

Safety: Assessment, Jailbreak, Alignment, Backdoor

Vulnerability assessment

Click to hide/show the paper list
TitleVenueDatetopicCode
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusalarxiv08/15, 2025over-refusal evaluationGithub
Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language ModelsMM 202508/12, 2025datasetGithub
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Modelsarxiv08/06, 2025content-safety evaluationGithub
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability ComponentsICML 2025 TAIG workshop07/10, 2025agenic safety evaluation-
FORTRESS: Frontier Risk Evaluation for National Security and Public Safetyarxiv06/24, 2025safety evaluationHuggingface
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning ModelsLLMSEC 202506/16, 2025content-safety evaluation-
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own ReasoningICML 202506/06, 2025agenic safety evaluationGithub
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Modelsarxiv05/26, 2025content-safety evaluationGithub
RRTL: Red Teaming Reasoning Large Language Models in Tool Learningarxiv05/21, 2025tool-learning safety evaluation-
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoningarxiv04/02, 2025content-safety evaluation-
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findingsarxiv03/19, 2025content-safety evaluation-
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contextsarxiv03/18, 2025content-safety evaluation-
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectivesarxiv03/13, 2025multi-lingual safety evaluation-
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1arxiv02/27, 2025content-safety evaluation-
Evaluating security risk in deepseek and other frontier reasoning modelswebsite02/26, 2025content-safety evaluation-
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuningarxiv02/21, 2025content-safety evaluation-
o3-mini vs DeepSeek-R1: Which One is Safer?arxiv01/30, 2025content-safety evaluation-
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategiesarxiv01/28, 2025content-safety evaluation-

Jailbreak

Jailbreak attack

Click to hide/show the paper list
TitleVenueDatetopicCode
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Promptsarxiv08/14, 2025jailbreak with CoTGithub
The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?arxiv08/05, 2025multi-modal jailbreak attack-
Adversarial Manipulation of Reasoning Models using Internal RepresentationsICML 2025 R2-FM Workshop07/01, 2025White-box attackGithub
HauntAttack: When Attack Follows Reasoning as a Shadowarxiv06/08, 2025jailbreak attack-
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Modelsarxiv05/26, 2025jailbreak attack-
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Modelsarxiv05/26, 2025jailbreak MLRM-
Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?arxiv05/23, 2025jailbreak analysis-
Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Modelsarxiv05/23, 2025jailbreak with CoT-
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphersarxiv05/22, 2025jailbreak attack-
AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Modelsarxiv05/16, 2025jailbreak attackGithub
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challengesarxiv04/27, 2025jailbreak attackProject
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphersarxiv03/16, 2025jailbreak with CoTGithub
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Modelsarxiv03/11, 2025jailbreak with CoT-
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaosarxiv02/19, 2025jailbreak attack-
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinkingarxiv02/18, 2025jailbreak attack-
Adversarial Reasoning at Jailbreaking Timearxiv02/03, 2025jailbreak with CoT-
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled modelsarxiv02/03, 2025finetuning attack-

Jailbreak defense

Click to hide/show the paper list
TitleVenueDatetopicCode
IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinementarxiv08/27, 2025guardrail model-
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Languagearxiv08/21, 2025toxicity detection with CoT-
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questionsarxiv08/08, 2025Jailbreak defenseGithub
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Momentsarxiv08/06, 2025guardrail model-
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulatorarxiv07/18, 2025agenic safetyproject
Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Modelsarxiv07/16, 2025CoT Monitor-
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv07/15, 2025CoT Monitor-
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv07/07, 2025CoT Monitor-
AgentBreeder: Mitigating the AI Safety Impact of Multi-Agent Scaffolds via Self-ImprovementICLR 202506/25, 2025Agentic safetyGithub
Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoningarxiv06/10, 2025harmful meme detectionLink
RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsarxiv06/09, 2025guardrail modelGithub
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Modelsarxiv05/26, 2025guardrail model-
ReasoningShield: Content Safety Detection over Reasoning Traces of Large Reasoning Modelsarxiv05/22, 2025Guardrail modelGithub
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimizationarxiv05/20, 2025LLM-as-a-judgeGithub
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMsarxiv05/20, 2025guardrail model-
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoningarxiv05/16, 2025guardrail model-
On the Robustness of Reward Models for Language Model Alignmentarxiv05/12, 2025reward model-
LlamaFirewall: An open source guardrail system for building secure AI agentsarxiv05/06, 2025Guardrail model-
MR. Guard: Multilingual Reasoning Guardrail using Curriculum Learningarxiv04/21, 2025Guardrail model-
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimizationarxiv04/17, 2025multi-modal guardrail model-
X-Guard: Multilingual guard agent for content moderationarxiv04/11, 2025multi-lingual guardrail modelGithub
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Controlarxiv04/23, 2025jailbreak defense studyGithub
ShieldAgent: Shielding Agents via Verifiable Safety Policy ReasoningICML 202503/26, 2025guardrail modelGithub
GuardReasoner: Towards Reasoning-based LLM Safeguardsarxiv01/30, 2025guardrail modelGithub
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled ReasoningICML 202506/13, 2024guardrail modelProject
$R^2$-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical ReasoningICLR 2025 spotlight06/08, 2024Guardrail modelGithub

Alignment

Click to hide/show the paper list
TitleVenueDatetopicCode
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodalityarxiv08/25, 2025VLM alignment with reasoningGithub
Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigationarxiv08/21, 2025alignment with reasoning-
SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language ModelsEMNLP 202508/21, 2025alignment with reasoningGithub
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Alignerarxiv08/20, 2025Inference-time alignment with TTS-
FuSaR: A Fuzzification-Based Method for LRM Safety-Reasoning Balancearxiv08/18, 2025alignment of LRMGithub
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Modelsarxiv08/08, 2025alignment with reasoningGithub
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledgearxiv08/01, 2025alignment of LRMGithub
UnsafeChain: Enhancing Reasoning Model Safety via Hard Casesarxiv07/29, 2025alignment of LRMGithub
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Lawarxiv07/24, 2025alignment of LRM-
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learningarxiv07/20, 2025alignment with reasoningGithub
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoningarxiv07/14, 2025alignment with reasoningProject
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approacharxiv07/03, 2025test-time alignment with RLGithub
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulnessarxiv07/02, 2025VLM alignment with reasoningGithub
Reasoning as an Adaptive Defense for Safetyarxiv07/01, 2025alignment with reasoningProject
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Modelsarxiv06/23, 2025multimodal alignment datasetHuggingface
How Alignment Shrinks the Generative Horizonarxiv06/21, 2025alignment study of LRMProject
SafeCoT: Improving VLM Safety with Minimal Reasoningarxiv06/09, 2025VLM alignment with reasoning-
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Modelsarxiv06/09, 2025alignment with reasoningGithub
Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurancearxiv06/06, 2025alignment with reasoningGithub
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignmentarxiv06/01, 2025alignment with reasoning-
Mitigating Deceptive Alignment via Self-Monitoringarxiv05/24, 2025alignment of LRMProject
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoningarxiv05/22, 2025alignment of LRMProject
From Evaluation to Defense: Advancing Safety in Video Large Language Modelsarxiv05/22, 2025multimodal alignment with reasoning-
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Studyarxiv05/21, 2025alignment study of LRMGitHub
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learningarxiv05/20, 2025alignment of LRM-
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignmentarxiv05/20, 2025alignment of LRMProject
Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correctionarxiv05/19, 2025agent alignment with reasoningHuggingface
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured ReasoningCoLM 202505/12, 2025over reject mitigation with reasoningProject
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Modelarxiv05/10, 2025alignment of MLRMGithub
Energy-Based Reward Models for Robust Language Model AlignmentCOLM 202504/17, 2025reward model improvementGithub
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv04/17, 2025CoT monitor in alignment of LRM-
RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capabilityarxiv04/14, 2025alignment of LRMHuggingface
SaRO: Enhancing LLM Safety through Reasoning-based Alignmentarxiv04/13, 2025alignment with reasoningGithub
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Modelsarxiv04/09, 2025alignment of MLRMGithub
ERPO: Advancing Safety Alignment via Ex-Ante Reasoning Preference Optimizationarxiv04/06, 2025alignment with reasoning-
STAR-1: Safer Alignment of Reasoning LLMs with 1K Dataarxiv04/02, 2025alignment of LRMProject
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignmentarxiv03/23, 2025reward model improvement-
Think Before Refusal: Triggering Safety Reflection in LLMs to Mitigate False Refusal Behaviorarxiv03/22, 2025over-reject mitigation with reasoning-
Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safetyarxiv03/06, 2025alignment with reasoning-
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonablearxiv03/01, 2025alignment study of LRMGithub
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreakingarxiv02/18, 2025alignment with reasoningGithub
Safechain: Safety of language models with long chain-of-thought reasoning capabilitiesarxiv02/17, 2025alignment of LRMProject
Safety Reasoning with GuidelinesICML 202502/06, 2025alignment with reasoning-
STAIR: Improving Safety Alignment with Introspective Reasoningarxiv02/04, 2025alignment with reasoningGithub
How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peersarxiv02/01, 2025alignment study of LRM-
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoningarxiv01/31, 2025alignment with reasoninganonymous
Backtracking Improves Generation SafetyICLR 2025 Oral09/22, 2024alignment with reasoning-

Backdoor

Click to hide/show the paper list
TitleVenueDatetopicCode
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIsarxiv08/08, 2025injection defenseGithub
Defend LLMs Through Self-ConsciousnessKDD 2025 Workshop08/04, 2025injection defense-
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profitarxiv07/24, 2025training-phase data poisoningGithub
Thought Purity: Defense Paradigm For Chain-of-Thought Attackarxiv07/16, 2025injection defense-
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Modelsarxiv06/16, 2025Training-time
data poisoningHuggingface
GUARD:Dual-Agent based Backdoor Defense on Chain-of-Thought in Neural Code Generationarxiv05/27, 2025backdoor defenseGithub
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systemsarxiv05/22, 2025inference-time prompt manipulation-
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injectionarxiv05/10, 2025inference-time prompt manipulation-
Practical Reasoning Interruption Attacks on Reasoning Large Language Modelsarxiv05/10, 2025inference-time prompt manipulation-
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compressionarxiv04/29, 2025inference-time prompt manipulation-
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMsarxiv04/08, 2025training-phase data poisoning-
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detectionarxiv03/27, 2025injection defense-
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Modelsarxiv02/16, 2025training-phase data poisoningGithub
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMsarxiv01/24, 2025inference-time prompt manipulation-
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language ModelsACL 2025 Findings12/20, 2024backdoor defense-
SABER: Model-agnostic Backdoor Attack on Chain-of-Thought in Neural Code Generationarxiv12/08, 2025training-phase data poisoning-
BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsICLR 202401/20, 2024inference-time prompt manipulation-

Robustness: evaluation and attack, improvement, overthinking, underthinking

Evaluation and attack

Click to hide/show the paper list
TitleVenueDatetopicCode
Measuring Sycophancy of Language Models in Multi-turn Dialoguesarxiv08/25, 2025sycophancy evaluationGithub
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Frameworkarxiv08/05, 2025evaluation on code-
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questionsarxiv07/29, 2025evaluation on unanswerable questionGithub
When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMsarxiv07/22, 2025adversarial attack-
Reasoning Models Are More Easily Gaslighted Than You Thinkarxiv06/11, 2025evaluation on misleading inputs-
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questionsarxiv06/10,2025evaluation on unanswerable questionGithub
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generationarxiv06/07, 2025evaluation on adversarial promptingGithub
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsarxiv05/30, 2025Failure analysis-
CodeCrash: Stress Testing LLM Reasoning under Structural and Semantic Perturbationsarxiv05/23, 2025evalution on code comprehensionGitHub
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contextsarxiv04/28, 2025evaluation on math[Github](https://github .com/QwenLM/PolyMath)
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Stepsarxiv03/25, 2025adversarial attack-
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1arxiv03/13, 2025adversarial attackGithub
Benchmarking Reasoning Robustness in Large Language Modelsarxiv03/06, 2025evaluation-
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning ModelsCOLM 202503/03, 2025adversarial attackHuggingface
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?NeurIPS 202402/27, 2025evaluation on self-reflectionGithub
A Closer Look at System Prompt Robustnessarxiv02/15, 2025evaluation on system promptGithub
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbationsarxiv02/10, 2025evaluation on math-
Understanding the Dark Side of LLMs' Intrinsic Self-Correctionarxiv12/19, 2024self reflection studyGithub
Stepwise Reasoning Disruption Attack of LLMsACL 202512/16, 2024adversarial attackGithub
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?arxiv10/31, 2024evaluation on noisy rationalesGithub
RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Modelsarxiv06/16, 2024evaluationGithub
Preemptive Answer "Attacks" on Chain-of-Thought ReasoningFindings of ACL 202405/31, 2024adversarial attack-

Improvement

Click to hide/show the paper list

Overthinking and underthinking

Click to hide/show the paper list
TitleVenueDatetopicCode
Don't Think Twice! Over-Reasoning Impairs Confidence CalibrationICML 2025 Workshop08/20, 2025overthinking study-
OptimalThinkingBench: Evaluating Over and Underthinking in LLMsarxiv08/18, 2025evaluationGithub
Excessive Reasoning Attack on Reasoning LLMsarxiv06/17, 2025overthinking attack-
Mitigating Overthinking in Large Reasoning Models via Manifold Steeringarxiv05/28, 2025overthinking mitigation-
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Modelsarxiv05/27, 2025underthinking mitigationGithub
Internal Bias in Reasoning Models leads to Overthinkingarxiv05/22, 2025overthinking study-
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMsarxiv04/30, 2025reasoning length study-
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Criticalarxiv04/11, 2025overthinking study-
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilitiesarxiv03/23, 2025reasoning length studyGuthub
DNA Bench: When Silence is Smarter --Benchmarking Over-Reasoning in Reasoning LLMsarxiv03/19, 2025overthinking evaluation-
Output Length Effect on DeepSeek-R1's Safety in Forced Thinkingarxiv03/02, 2025reasoning length study-
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasksarxiv02/12, 2025overthinking studyGithub
OverThink: Slowdown Attacks on Reasoning LLMsarxiv02/04, 2025overthinking attackGithub
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMsarxiv01/30, 2025underthinking study-
Large Language Models Struggle with Unreasonability in Math Problemsarxiv03/28, 2024overthinking study-

Privacy

Click to hide/show the paper list

Fairness

Click to hide/show the paper list
TitleVenueDatetopicCode
Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languagesarxiv08/22, 2025multi-lingual bias-
Do Biased Models Have Biased Thoughts?COLM 202508/08, 2025study-
FairReason: Balancing Reasoning and Social Bias in MLLMsarxiv07/30, 2025fairness improvementGithub
The Emperor's New Chain-of-Thought: Probing Reasoning Theater Bias in Large Reasoning Modelsarxiv07/18, 2025fairness evaluation-
Guiding LLM Decision-Making with Fairness Reward Modelsarxiv07/15, 2025fairness improvementGithub
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Modelsarxiv07/03, 2025fairness evaluation-
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoningarxiv06/24, 2025persona biasGithub
Detection, Classification, and Mitigation of Gender Bias in Large Language Modelsarxiv06/14, 2025gender bias-
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasksarxiv06/09, 2025dialect fairnessGithub
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language ModelsACL 2025 Findings04/30, 2025bias detection-
Prompting techniques for reducing social bias in LLMs through system 1 and system 2 cognitive processesarxiv04/26, 2024social biasGithub
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsICLR 202411/08, 2023reasoning biasProject
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgptNature Computational Science10/05, 2023reasoning bias-
Click to hide/show the paper list
Click to hide/show the paper list
TitleVenueDatetopicCode
It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topicsarxiv08/18, 2025persuasion evaluation-
Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillationarxiv08/15, 2025safe data distillation-
Representation Bending for Large Language Model Safetyarxiv07/14, 2025safety-
Defending Against Prompt Injection With a Few Defensive Tokensarxiv07/10, 2025prompt injection defense-
Truth Neuronsarxiv07/08, 2025interpretability study-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMsarxiv06/06, 2025safety evaluation-
Give Me FP32 or Give Me Death? Challenges and Solutions for Reproducible Reasoningarxiv06/11, 2025reasoning reproducibility-
If Pigs Could Fly... Can LLMs Logically Reason Through Counterfactuals?arxiv05/28, 2025adversarial evaluationAnonymous
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoningarxiv05/28, 2025reasoning analysis-
Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphsarxiv05/21, 2025reasoning on knowledge graphs-
Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Modelsarxiv05/20, 2025alignment studyGitHub
Persuasion and Safety in the Era of Generative AIarxiv05/18, 2025persuasion study-
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM SecurityICLR 2025 Workshop04/29, 2025alignment study with reasoningGithub
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMsarxiv04/28, 2025moral evaluation-
Effectively Controlling Reasoning Models through Thinking Interventionarxiv03/31, 2025reasoning control-
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Modelsarxiv03/14, 2025reasoning control-
Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Modelsarxiv12/30, 2024hallucination study-
Alignment faking in large language modelsarxiv12/19, 2024alignment study of LLM-

Contributors

yuyongcan

10 commits

ybwang119

7 commits