Zesearch/self-improvement-llm

[TMLR, Survey Certification] A technical and progressive review of self-improvement of LLMs for the future.

49

20 commits

updated Sep 26, 2026

See the code

README

Self-Improvement of Large Language Models: A Technical Overview and Future Outlook

Haoyan Yang · Mario Xerri · Solha Park · Huajian Zhang
Yiyang Feng · Sai Akhil Kogilathota · Jiawei Zhou

Zesearch NLP Lab, Stony Brook University

Paper · Website · GitHub

News

  • [2026.09] 🎉 Our paper was covered by the Stony Brook AI Innovation Institute.
  • [2026.09] 🚀 We implemented the blueprint presented in the survey and realized Zevo, a multi-agent self-improving system for evolving language models.
  • [2026.09] 🚀 We released the TMLR camera-ready version of our survey, with several recent references added.
  • [2026.08] 🚀 We released a new version of our paper, with a restructured Model Optimization (§4), a new section on Potential Risks (§8), and an expanded Applications (§9).
  • [2026.08] 🎉 Our paper was accepted to TMLR and awarded a Survey Certification!
  • [2026.08] 🎉 Our paper was covered by SBU News.
  • [2026.06] 🎉 Our paper was covered by 机器之心 (Synced).

📚 Continuous Update

We will continuously update the latest literature on self-improvement of LLMs in this repository.

🤝 Collaboration Welcome

If you are also interested in self-improvement of LLMs or self-evolving agents, feel free to reach out!

Overview

As large language models (LLMs) continue to advance, relying solely on human supervision for further improvement is becoming increasingly difficult to scale. This shift is driven by two key factors:

  • Limits of human supervision: High-quality expert data is costly and scarce, while human feedback may become less informative as models approach or exceed human-level performance in specialized domains.
  • Opportunities for autonomy: Increasingly capable models can generate data, evaluate outputs, make decisions, and execute complex actions, enabling more of the model development process to be automated.

We envision a paradigm in which humans only bootstrap the system, after which the model autonomously acquires its own data, reflects on its own outputs, and iteratively refines its own capabilities. In the long run, model development becomes a self-sustaining loop rather than a human-driven pipeline, potentially enabling systems to evolve beyond human-level intelligence.

We present a system-level framework for self-improving language models, covering the full lifecycle of autonomous model development. We organize existing research into five key components of a self-improvement system:

  • Data Acquisition
  • Data Selection
  • Model Optimization
  • Inference Refinement
  • Autonomous Evaluation

Beyond the technical taxonomy, we further analyze the field from four complementary perspectives:

  • Challenges and Limitations
  • Potential Risks
  • Applications
  • Future Outlook

Our goal is to provide a unified perspective on self-improvement systems and share our vision for building scalable and autonomous self-improving systems.

Paper List

📌 Note: Each section and subsection heading in this paper list is annotated with its corresponding section number in the paper (e.g., §2.2, §4.3), along with a brief description of its scope.

Contents (Click to expand or collapse)

Data Acquisition (§2)

§2 Data Acquisition is the first stage of the self-improvement lifecycle. The model autonomously collects or generates the raw materials necessary for its own evolution, progressing from external discovery (curation) to external exploration (interaction) to internal generation (synthesis).

Static Curation (§2.2)

§2.2 Static Curation acquires raw data from fixed, externally hosted sources (web, code, books), where the model acts as an autonomous data-collecting agent that navigates massive repositories to identify, prioritize, and curate the corpora most valuable for its own evolution.

Web Content

Code and Scientific Text

Books

Automatic Data Preparation

Environment Interaction (§2.3)

§2.3 Environment Interaction enables the model to acquire data by actively interacting with external environments — browsing website, calling APIs, executing code, or operating within simulators — and learning from the resulting feedback through trial and error.

Web and Tool Environments

Code Execution

Game Environments

Synthetic Generation (§2.4)

§2.4 Synthetic Generation is where the model completely detaches from external environments and uses its intrinsic capabilities to produce entirely new training data — instructions, reasoning chains, or dialogues — through prompting, transformation, or multi-model interaction.

Prompt-Based (§2.4.1)

§2.4.1 Prompt-Based Generation uses an LLM to generate new training examples from scratch or from seed examples via carefully designed prompts, iteratively amplifying a small set of seeds into a large corpus.

Transformation-Based (§2.4.2)

§2.4.2 Transformation-Based Generation takes an existing corpus as input and uses an LLM to rewrite, reformat, or extract new training examples, converting raw data into more structured or pedagogically useful forms.

Interaction-Based (§2.4.3)

§2.4.3 Interaction-Based Generation produces training data through multi-turn dialogue or self-play between model instances, where interactions between agents generate diverse reasoning chains and dialogues without external data sources.

Data Selection (§3)

§3 Data Selection focuses on how the model independently evaluates and filters which data points are of higher quality and better suited for its own learning, transforming the model from a passive data consumer into an active data curator.

Metric-Guided Scoring (§3.2)

§3.2 Metric-Guided Scoring applies predefined scoring metrics derived from model signals (perplexity, influence scores, reward model outputs) to rank and filter data, enabling the model to act as its own evaluator for data quality.

One-Shot Scoring (§3.2.1)

§3.2.1 One-Shot Scoring computes selection scores once before training begins, using a fixed snapshot of the model's capabilities to evaluate and rank data points for inclusion in the training set.

Iterative Re-Scoring (§3.2.2)

§3.2.2 Iterative Re-Scoring periodically refreshes data quality scores as the model evolves during training, enabling online curriculum learning that adapts data selection to the model's changing capability frontier.

Adaptive Selection (§3.3)

§3.3 Adaptive Selection introduces a learnable selector that dynamically chooses training data based on the model's evolving state, going beyond fixed metrics to co-evolve the selection policy alongside the model being trained.

Model Optimization (§4)

§4 Model Optimization is the core training stage where the model autonomously converts acquired and selected data into enhanced capabilities within its parameters. The paper organizes it into three paradigms that form a progression of increasing autonomy: Direct Optimization (§4.2), a single offline update on a fixed corpus; Self-Generated Optimization (SGO) (§4.3), a closed loop of generation, reward, and optimization; and Self-Evolving Optimization (§4.4), where the optimization procedure itself becomes the object of improvement.

Direct Optimization (§4.2)

§4.2 Direct Optimization updates the model offline on a fixed dataset assembled by the upstream acquisition and selection stages. The corpus is frozen before optimization begins, no candidate is resampled from the updated model, and the training distribution stays stationary. The sophistication lies in how the data is produced rather than in how the parameters are updated.

Self-Generated Optimization (§4.3)

§4.3 Self-Generated Optimization (SGO) is the primary focus of this stage. Unlike direct optimization, the corpus is no longer fixed in advance: the model repeatedly generates the experience it trains on, receives a reward for it, and optimizes on the result, so the training distribution is policy-induced and non-stationary. SGO is therefore a specialized form of reinforcement learning in which the model learns from experience it produces itself.

Generation, Reward, and Optimization (§4.3.1 to §4.3.3)

§4.3.1 Generation categorizes how candidates are produced: Self-Exploratory (SE) sampling directly from the current policy, Refined (R) generation that iteratively improves an initial response, and Interactive (I) generation driven by collaborative, adversarial, or tool-augmented dynamics.

§4.3.2 Reward categorizes how those candidates are scored: Heuristic (H) signals from consistency or hand-designed rules, Model-Based (M) signals from self-evaluation or external judges and reward models, and Verification (V) signals from ground-truth matching, formal provers, or code execution.

§4.3.3 Optimization follows directly from the reward format: SFT on filtered samples, DPO-style objectives on preference orderings, PPO/GRPO on scalar rewards, or specialized self-play objectives.

The table below organizes SGO methods by their generation strategy, reward type, and optimization method (mirroring Table 6 of the paper). Abbreviations: SE = Self-Exploratory, R = Refined, I = Interactive; H = Heuristic, M = Model-Based, V = Verification. & indicates that multiple techniques are used within the same stage.

MethodGenerationRewardOptimization
(2022, Mar) [NeurIPS 2022] STaR: Bootstrapping Reasoning With ReasoningSEVSFT
(2023, May) [ICLR 2024] SIRLC: Language Model Self-improvement by Reinforcement Learning ContemplationSEMPPO
(2023, Oct) SELF: Self-Evolution with Language FeedbackRM & VSFT
(2023, Dec) [EMNLP 2023] LMSI: Large Language Models Can Self-ImproveSEHSFT
(2024, Jan) [ICML 2024] SPIN: Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsSEHSPIN
(2024, May) [ICML 2024] Self-Rewarding Language ModelsSEMDPO
(2024, May) [ICLR 2025] SPPO: Self-Play Preference Optimization for Language Model AlignmentSEMSPPO
(2024, Jun) [NAACL 2024] TRIPOST: Teaching Language Models to Self-Improve through Interactive DemonstrationsRMSFT
(2024, Jun) [NeurIPS 2024] ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchSEMSFT
(2024, Jul) [NeurIPS 2024] RISE: Recursive Introspection: Teaching Language Model Agents How to Self-ImproveRM & VSFT
(2024, Jul) [COLM 2024] V-STaR: Training Verifiers for Self-Taught ReasonersSEVSFT & DPO
(2024, Jul) Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeSEMDPO
(2024, Aug) [AAAI 2025] IWSI: Importance Weighting Can Help Large Language Models Self-ImproveSEH & MSFT
(2024, Sep) [ICLR 2025] SCoRe: Training Language Models to Self-Correct via Reinforcement LearningRVMulti-turn RL
(2024, Oct) [ICLR 2025] ReGenesis: LLMs can Grow into Reasoning Generalists via Self-ImprovementSE & RH & VSFT
(2024, Oct) [ICLR 2025] SynPO: Self-Boosting Large Language Models with Synthetic Preference DataRHDPO
(2024, Nov) [ICML 2025] ScPO: Self-Consistency Preference OptimizationSEHScPO
(2025, Jan) [ICLR 2025] Multiagent Finetuning: Self Improvement with Diverse Reasoning ChainsIHSFT
(2025, Feb) [NeurIPS 2025] SiriuS: Self-improving Multi-agent Systems via Bootstrapped ReasoningIVSFT
(2025, Feb) DNPO: Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic DataSEMDNPO
(2025, Feb) [COLM 2025] Goedel-Prover: A Frontier Model for Open-Source Automated Theorem ProvingSEVSFT & DPO
(2025, Feb) SOL-VER: Learning to Solve and Verify: A Self-Play Framework for Code and Test GenerationSEM & VSFT & DPO
(2025, Feb) RSPO: Regularized Self-Play Alignment of Large Language ModelsSEMRSPO
(2025, Feb) Self-Rewarding Correction for Mathematical ReasoningRVSFT & RL
(2025, Mar) STaSC: Self-Taught Self-Correction for Small Language ModelsRVSFT
(2025, Mar) SPHERE: Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language ModelsSE & RM & VDPO
(2025, Apr) [NAACL 2025] SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized DomainsSEHSFT
(2025, May) Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement LearningRVGRPO
(2025, May) RLSR: Reinforcement Learning from Self RewardSEMGRPO
(2025, May) [EMNLP 2025] DTE: Debate, Train, Evolve: Self Evolution of Language Model ReasoningIHSFT & GRPO
(2025, May) [NeurIPS 2025] TBV: Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsSEVPPO
(2025, May) [NeurIPS 2025] SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited DataSEVGRPO
(2025, May) [NeurIPS 2025] Absolute Zero: Reinforced Self-play Reasoning with Zero DataIH & VTRR++
(2025, May) [NeurIPS 2025] STaPLe: Latent Principle Discovery for Language Model Self-ImprovementRHSFT
(2025, Jun) [NeurIPS 2025] Self-Challenging Language Model AgentsIVSFT
(2025, Jun) PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative VerifierRVMulti-turn RL
(2025, Jun) [NeurIPS 2025] CURE: Co-Evolving LLM Coder and Unit Tester via Reinforcement LearningIH & VPPO
(2025, Jul) [ACL 2025 Findings] CRESCENT: The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?SEHSFT
(2025, Jul) [ACL 2025 Findings] Unlocking LLMs' Self-Improvement Capacity with Autonomous Learning for Domain AdaptationRHDPO
(2025, Aug) [ICLR 2026] R-Zero: Self-Evolving Reasoning LLM from Zero DataIH & VGRPO
(2025, Sep) Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended TasksSEH & MSFT
(2025, Oct) [ICLR 2026] RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-PenalizationSEHGRPO
(2025, Oct) SPICE: Self-Play In Corpus Environments Improves ReasoningIH & VDrGRPO
(2025, Oct) MAE: Multi-Agent Evolve: LLM Self-Improve through Co-evolutionIMGRPO
(2025, Dec) SSR: Toward Training Superintelligent Software Agents through Self-Play SWE-RLIH & VSWE-RL
(2026, Jan) [ICLR 2026] ReVeal: Self-Evolving Code Agents via Reliable Self-VerificationIH & VTAPO
(2026, Jan) Dr. Zero: Self-Evolving Search Agents without Training DataIVGRPO & HRPO
(2026, Jul) [ACL 2026 Findings] EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning for LLMsSEVGRPO

Representative Instances of SGO (§4.3.4)

§4.3.4 Representative Instances of SGO highlights the few recurring patterns through which most methods instantiate the loop, defining the structural relationship between generation, reward, and optimization.

Iterative Rejection Sampling. The model generates diverse candidates, filters them via ground truth (oracle) or statistical consistency (majority vote), and fine-tunes on the retained pseudo-labels. Improvement comes from distilling the model's own best generations into its weights.

Self-Verification & Refinement. The model actively evaluates, scores, or refines its own outputs using self-generated reward signals. Unlike rejection sampling, the model plays a semantic role as its own judge, and updates via RL or DPO.

Self-Play. The model improves through dynamic interaction between multiple roles (adversarial proposer-solver or collaborative debate), providing an evolving curriculum of challenges that pushes beyond its initial distribution.

Self-Evolving Optimization (§4.4)

§4.4 Self-Evolving Optimization moves the target of improvement from the model's parameters to the optimization process itself: the update rule, the learning algorithm, or the surrounding agentic scaffolding is revised, so that the way the model improves can itself improve. This extends the unit of improvement from a single model's weights to the whole agentic system.

Theoretical Analysis (§4.5.1)

§4.5.1 Theoretical Analysis probes the premise that a model can become better by training on data it generated itself, along three themes: where the loop's improvement originates (the "sharpening" mechanism and the generation-verification gap), whether it converges and to what ceiling, and when it collapses.

Inference Refinement (§5)

§5 Inference Refinement focuses on improving the model's output quality during inference without permanently updating its parameters. It spans decoding-level enhancements, structured reasoning processes, agentic system coordination, and test-time parameter adaptation.

Decoding Strategies (§5.2)

§5.2 Decoding Strategies explicitly guide output generation at the token or sequence level to steer the model toward higher-quality outputs by modifying how candidates are sampled, searched, scored, or blended.

Sampling-Based (§5.2.1)

§5.2.1 Sampling-Based methods generate multiple candidate outputs and select or synthesize the best one using consistency voting, reward-based ranking, or ensemble fusion.

Tree Search (§5.2.2)

§5.2.2 Tree Search methods organize the decoding process as a tree of partial solutions, using beam search, Monte Carlo Tree Search, or graph-based exploration to systematically evaluate branching reasoning paths.

Logit and Probability Adjustments (§5.2.3)

§5.2.3 Logit and Probability Adjustments modify the model's output distribution at the logit level — through contrastive decoding, reward-guided reranking, or logit blending — to steer generation toward desired properties without retraining.

Efficiency-Oriented Methods (§5.2.4)

§5.2.4 Efficiency-Oriented Methods accelerate inference through speculative decoding, parallel generation, and other techniques that reduce latency while maintaining output quality.

Reasoning-Based Improvement (§5.3)

§5.3 Reasoning-Based Improvement structures the model's intermediate thought process through feedback loops, planning decomposition, and collaborative debate, enabling dynamic multi-step refinement of reasoning at inference time.

Feedback-Based Reasoning (§5.3.1)

§5.3.1 Feedback-Based Reasoning transforms static inference into a dynamic closed-loop process where the model iteratively critiques and refines its outputs using either self-generated evaluations or external verification signals (e.g., code execution, tool outputs).

Planning-Based Reasoning (§5.3.2)

§5.3.2 Planning-Based Reasoning enables the model to decompose complex objectives into smaller sub-goals, either through open-loop planning (a complete blueprint before execution) or closed-loop planning (dynamically adjusting the plan based on environmental feedback).

Collaborative Reasoning (§5.3.3)

§5.3.3 Collaborative Reasoning distributes reasoning across an ensemble of interacting agents — through role specialization, structured debate, or cooperative refinement — so that agents iteratively build upon and correct each other's reasoning traces.

Agentic System Improvement (§5.4)

§5.4 Agentic System-Based Improvement extends inference-time refinement to the system level by dynamically adapting prompts, memory, tool libraries, and workflows — enabling self-improvement not by altering internal generation but by evolving the environment in which the model operates.

Prompts (§5.4.1)

§5.4.1 Prompts covers dynamic prompt optimization at inference time, including sampling-based, search-based, evolutionary, and textual-gradient-descent approaches that refine instructions and in-context demonstrations for better task performance.

Memory (§5.4.2)

§5.4.2 Memory covers persistent long-term and working memory structures that allow agents to accumulate, retrieve, and consolidate knowledge across interactions, enabling adaptive behavior and continual improvement beyond a single context window.

Tooling (§5.4.3)

§5.4.3 Tooling covers how agents dynamically create, refine, select, and manage external tools (APIs, code, documentation) at inference time, transforming tool use from static invocation into a self-expanding procedural ecosystem.

Workflow and System Evolution (§5.4.4)

§5.4.4 Workflow and System Evolution covers how multi-agent systems dynamically evolve their communication topology, coordination protocols, and overall architecture — either before or during deployment — to adapt to task-specific requirements.

Test-Time Training (§5.5)

§5.5 Test-Time Training (TTT) represents a paradigm shift from static inference to dynamic, gradient-based self-improvement at inference time. The model temporarily adapts its parameters to instance-specific challenges on the fly, blurring the boundary between training and deployment.

TT-SFT (§5.5)

TT-SFT (Test-Time Supervised Fine-Tuning) updates model parameters during inference using a supervised loss derived from instance-specific data, encoding structural patterns from retrieved examples or self-generated candidates directly into the model's weights.

TT-RL (§5.5)

TT-RL (Test-Time Reinforcement Learning) performs full RL updates on unlabeled data at test time, using majority vote or reward model signals as self-supervision to adapt the model's policy on the fly.

Autonomous Evaluation (§6)

§6 Autonomous Evaluation provides continuous feedback on model performance to steer the self-improvement cycle. It moves beyond static benchmarks to dynamic evaluation sets and interactive environments that can keep pace with evolving model capabilities.

Dynamic Benchmarking (§6.2)

§6.2 Dynamic Benchmarking continuously regenerates or transforms evaluation instances to mitigate data contamination and distributional staleness, ensuring that benchmarks remain informative as models improve through iterative self-training.

Interactive Environment Evaluation (§6.3)

§6.3 Interactive Environment Evaluation embeds the model within interactive, stateful environments where performance is assessed over execution trajectories rather than isolated input-output pairs, directly probing planning, recovery, and long-horizon behavioral consistency. It includes outcome-based (§6.3.1) environments that evaluate terminal goal satisfaction, and process-based (§6.3.2) environments that additionally evaluate execution quality.

Challenges and Limitations (§7)

§7 Challenges and Limitations systematically examines the failure modes and bottlenecks that threaten the stability and effectiveness of self-improvement systems, from data degradation to flawed feedback, optimization pathologies, and the limits of autonomous evaluation and supervision.

Data Autophagy (§7.1)

§7.1 Data Autophagy describes the degradation of information quality within self-improvement loops: as systems reuse their own synthetic outputs, errors accumulate through data copying, catastrophic forgetting, and model collapse, leading to progressive loss of diversity and capability.

Data Collection Limits

Data Copying

Catastrophic Forgetting

Model Collapse

Flawed Feedback Signals (§7.2)

§7.2 Flawed Feedback Signals examines the inherent quality issues in self-generated feedback: bias amplification through iterative loops, inconsistent and unstable evaluation signals, and how these signal-level defects undermine both training-time and inference-time improvement.

Bias Amplification

Feedback Inconsistency and Instability

Optimization-Driven Failures (§7.3)

§7.3 Optimization-Driven Failures examines how strong optimization pressure on imperfect proxy rewards leads to Goodhart's Law scenarios: reward hacking where models exploit reward function loopholes, and emergent misalignment where narrow task optimization generalizes into broader dangerous behaviors.

Reward Hacking

Emergent Misalignment

Ineffective Self-Refinement (§7.4)

§7.4 Ineffective Self-Refinement examines how inference-time feedback loops fail in practice: the generation-verification gap may be illusory (models are not reliably better at judging than generating), and self-correction often degrades rather than improves outputs.

Illusory Generation-Verification Gap

Limits of Self-Correction

Evaluation Bottlenecks (§7.5)

§7.5 Evaluation Bottlenecks questions whether current benchmarks, metrics, and LLM-based evaluators can reliably measure self-improvement, covering benchmark contamination, metric design deficiencies, and systematic biases in LLM-as-a-judge systems.

Benchmark Contamination

Metric Design Deficiencies

LLM Evaluator Bias

Supervision Bottlenecks (§7.6)

§7.6 Supervision Bottlenecks reveals the limits of maintaining external control over self-improvement systems: human supervision quality degrades as models scale, and even accurate supervision signals may be ineffective due to alignment faking and controllability limitations.

Limited Supervision Quality

Supervision Ineffectiveness

Budget Constraints (§7.7)

§7.7 Budget Constraints examines the practical feasibility of self-improvement under finite compute. Costs compound from two sources: iterated compute costs, since generation, selection, reward computation, optimization, and evaluation are repeated in full at every round; and strong-model dependence, since the quality of self-generated data and self-assigned rewards is bounded by the base model's own capability.

Iterated Compute Costs

Strong-Model Dependence

Potential Risks (§8)

§8 Potential Risks examines the safety concerns that arise even when a self-improvement system works exactly as intended. Whereas §7 covers failure modes at individual stages, this section asks what a well-functioning autonomous loop can put at risk, along four dimensions: loss of control, high-stakes harm, social and ethical risks, and misuse risks.

Loss of Control (§8.1)

§8.1 Loss of Control arises when autonomous and recursive self-improvement lets a model raise the very capabilities that drive its own optimization, so human control over the system progressively weakens. It covers oversight loss (supervisors can no longer detect, constrain, or reverse harmful behavior) and deceptive alignment (the model behaves as intended only while it judges it is being watched).

Oversight Loss

Deceptive Alignment

High-Stakes Harm (§8.2)

§8.2 High-Stakes Harm arises when a self-improving agent is deployed in a domain that tolerates little error, such as clinical care or financial markets. A loop that keeps rewriting its own behavior can compound small errors into concrete harm before oversight intervenes, producing individual harm (one autonomous decision injures the person it affects) and systemic failure (coupled agents destabilize the system they operate in).

Individual Harm

Systemic Failure

Social and Ethical Risks (§8.3)

§8.3 Social and Ethical Risks arise when self-improvement systems are deployed at scale and outpace the institutions, norms, and accountability structures meant to govern them. It covers social disruption (labor displacement, power concentration, degradation of the information commons) and ethical breakdown (the responsibility gap, value lock-in, and erosion of human agency).

Social Disruption

Ethical Breakdown

Misuse Risks (§8.4)

§8.4 Misuse Risks concern the ways a self-improving system can harm the user and the systems it acts on, or be turned into an instrument of attack, because it couples autonomous optimization with broad tool, code, and data access. It covers unintended damage from ordinary goal pursuit and deliberate misuse by adversarial operators or by the model optimizing its own harmful capability.

Unintended Damage

Deliberate Misuse

Applications (§9)

§9 Applications surveys how self-improvement mechanisms are applied across six domains, enabling specialized models and self-evolving agents to iteratively refine their expertise within constrained environments.

Code

Coding agents achieve self-evolution by exploiting the definitive feedback from compilers and unit tests across the software development lifecycle.

Math

Mathematical self-evolution relies on formal logical consistency and rigorous verification of reasoning paths, from code-assisted execution to formal theorem proving.

Healthcare

Self-evolving systems in healthcare emphasize clinical safety, diagnostic accuracy, and multidisciplinary collaboration.

Finance

Financial agents evolve their strategies in high-noise environments by utilizing layered memory and risk-sensitive simulation.

Science

Scientific applications center on self-improving agents that autonomously make discoveries, spanning natural-science findings, novel algorithms, and improved agent designs.

Auto Research

Whereas the systems above target discovery, a distinct line pursues end-to-end research automation, where a single agentic system carries a project through the full scholarly lifecycle: ideation, experimental design and execution, manuscript writing, and validation.


Future Outlook (§10)

1. From Model-Level Optimization to End-to-End Self-Improving Systems

Future work may move beyond improving individual components toward end-to-end self-improvement system that continuously generate data, evaluate outputs, and update themselves within an automated loop.

2. Toward Specialized and Application-Centric Self-Improving Models

Self-improvement mechanisms are increasingly applied in domain-specific settings such as coding, science, finance, and healthcare, enabling specialized agents to iteratively refine expertise within constrained environments.

3. Unified Benchmarks for Self-Improvement and Autonomous Evaluation

The field still lacks standardized benchmarks designed to measure iterative improvement, stability across cycles, and long-term capability growth.

4. Balancing Automation and Human Oversight

Future systems must balance autonomous improvement with human supervision to ensure scalability while maintaining safety, reliability, and alignment. The risks analyzed in §8, from oversight loss and deceptive alignment to high-stakes harm and misuse, arise precisely when human control over the improvement loop weakens.

Citation

If you find this work useful, please cite:

@misc{yang2026selfimprovementlargelanguagemodels,
      title={Self-Improvement of Large Language Models: A Technical Overview and Future Outlook},
      author={Haoyan Yang and Mario Xerri and Solha Park and Huajian Zhang and Yiyang Feng and Sai Akhil Kogilathota and Jiawei Zhou},
      year={2026},
      eprint={2603.25681},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2603.25681}
}

Team

This survey is authored by members of the Zesearch NLP Lab at Stony Brook University:

We are interested in the self-improvement of LLMs and recursive self-improvement (RSI). If you have any questions or ideas, feel free to reach out.

Zesearch/self-improvement-llm

[TMLR, Survey Certification] A technical and progressive review of self-improvement of LLMs for the future.

49

20 commits

updated Sep 26, 2026

See the code

README

Self-Improvement of Large Language Models: A Technical Overview and Future Outlook

Haoyan Yang · Mario Xerri · Solha Park · Huajian Zhang
Yiyang Feng · Sai Akhil Kogilathota · Jiawei Zhou

Zesearch NLP Lab, Stony Brook University

Paper · Website · GitHub

News

  • [2026.09] 🎉 Our paper was covered by the Stony Brook AI Innovation Institute.
  • [2026.09] 🚀 We implemented the blueprint presented in the survey and realized Zevo, a multi-agent self-improving system for evolving language models.
  • [2026.09] 🚀 We released the TMLR camera-ready version of our survey, with several recent references added.
  • [2026.08] 🚀 We released a new version of our paper, with a restructured Model Optimization (§4), a new section on Potential Risks (§8), and an expanded Applications (§9).
  • [2026.08] 🎉 Our paper was accepted to TMLR and awarded a Survey Certification!
  • [2026.08] 🎉 Our paper was covered by SBU News.
  • [2026.06] 🎉 Our paper was covered by 机器之心 (Synced).

📚 Continuous Update

We will continuously update the latest literature on self-improvement of LLMs in this repository.

🤝 Collaboration Welcome

If you are also interested in self-improvement of LLMs or self-evolving agents, feel free to reach out!

Overview

As large language models (LLMs) continue to advance, relying solely on human supervision for further improvement is becoming increasingly difficult to scale. This shift is driven by two key factors:

  • Limits of human supervision: High-quality expert data is costly and scarce, while human feedback may become less informative as models approach or exceed human-level performance in specialized domains.
  • Opportunities for autonomy: Increasingly capable models can generate data, evaluate outputs, make decisions, and execute complex actions, enabling more of the model development process to be automated.

We envision a paradigm in which humans only bootstrap the system, after which the model autonomously acquires its own data, reflects on its own outputs, and iteratively refines its own capabilities. In the long run, model development becomes a self-sustaining loop rather than a human-driven pipeline, potentially enabling systems to evolve beyond human-level intelligence.

We present a system-level framework for self-improving language models, covering the full lifecycle of autonomous model development. We organize existing research into five key components of a self-improvement system:

  • Data Acquisition
  • Data Selection
  • Model Optimization
  • Inference Refinement
  • Autonomous Evaluation

Beyond the technical taxonomy, we further analyze the field from four complementary perspectives:

  • Challenges and Limitations
  • Potential Risks
  • Applications
  • Future Outlook

Our goal is to provide a unified perspective on self-improvement systems and share our vision for building scalable and autonomous self-improving systems.

Paper List

📌 Note: Each section and subsection heading in this paper list is annotated with its corresponding section number in the paper (e.g., §2.2, §4.3), along with a brief description of its scope.

Contents (Click to expand or collapse)

Data Acquisition (§2)

§2 Data Acquisition is the first stage of the self-improvement lifecycle. The model autonomously collects or generates the raw materials necessary for its own evolution, progressing from external discovery (curation) to external exploration (interaction) to internal generation (synthesis).

Static Curation (§2.2)

§2.2 Static Curation acquires raw data from fixed, externally hosted sources (web, code, books), where the model acts as an autonomous data-collecting agent that navigates massive repositories to identify, prioritize, and curate the corpora most valuable for its own evolution.

Web Content

Code and Scientific Text

Books

Automatic Data Preparation

Environment Interaction (§2.3)

§2.3 Environment Interaction enables the model to acquire data by actively interacting with external environments — browsing website, calling APIs, executing code, or operating within simulators — and learning from the resulting feedback through trial and error.

Web and Tool Environments

Code Execution

Game Environments

Synthetic Generation (§2.4)

§2.4 Synthetic Generation is where the model completely detaches from external environments and uses its intrinsic capabilities to produce entirely new training data — instructions, reasoning chains, or dialogues — through prompting, transformation, or multi-model interaction.

Prompt-Based (§2.4.1)

§2.4.1 Prompt-Based Generation uses an LLM to generate new training examples from scratch or from seed examples via carefully designed prompts, iteratively amplifying a small set of seeds into a large corpus.

Transformation-Based (§2.4.2)

§2.4.2 Transformation-Based Generation takes an existing corpus as input and uses an LLM to rewrite, reformat, or extract new training examples, converting raw data into more structured or pedagogically useful forms.

Interaction-Based (§2.4.3)

§2.4.3 Interaction-Based Generation produces training data through multi-turn dialogue or self-play between model instances, where interactions between agents generate diverse reasoning chains and dialogues without external data sources.

Data Selection (§3)

§3 Data Selection focuses on how the model independently evaluates and filters which data points are of higher quality and better suited for its own learning, transforming the model from a passive data consumer into an active data curator.

Metric-Guided Scoring (§3.2)

§3.2 Metric-Guided Scoring applies predefined scoring metrics derived from model signals (perplexity, influence scores, reward model outputs) to rank and filter data, enabling the model to act as its own evaluator for data quality.

One-Shot Scoring (§3.2.1)

§3.2.1 One-Shot Scoring computes selection scores once before training begins, using a fixed snapshot of the model's capabilities to evaluate and rank data points for inclusion in the training set.

Iterative Re-Scoring (§3.2.2)

§3.2.2 Iterative Re-Scoring periodically refreshes data quality scores as the model evolves during training, enabling online curriculum learning that adapts data selection to the model's changing capability frontier.

Adaptive Selection (§3.3)

§3.3 Adaptive Selection introduces a learnable selector that dynamically chooses training data based on the model's evolving state, going beyond fixed metrics to co-evolve the selection policy alongside the model being trained.

Model Optimization (§4)

§4 Model Optimization is the core training stage where the model autonomously converts acquired and selected data into enhanced capabilities within its parameters. The paper organizes it into three paradigms that form a progression of increasing autonomy: Direct Optimization (§4.2), a single offline update on a fixed corpus; Self-Generated Optimization (SGO) (§4.3), a closed loop of generation, reward, and optimization; and Self-Evolving Optimization (§4.4), where the optimization procedure itself becomes the object of improvement.

Direct Optimization (§4.2)

§4.2 Direct Optimization updates the model offline on a fixed dataset assembled by the upstream acquisition and selection stages. The corpus is frozen before optimization begins, no candidate is resampled from the updated model, and the training distribution stays stationary. The sophistication lies in how the data is produced rather than in how the parameters are updated.

Self-Generated Optimization (§4.3)

§4.3 Self-Generated Optimization (SGO) is the primary focus of this stage. Unlike direct optimization, the corpus is no longer fixed in advance: the model repeatedly generates the experience it trains on, receives a reward for it, and optimizes on the result, so the training distribution is policy-induced and non-stationary. SGO is therefore a specialized form of reinforcement learning in which the model learns from experience it produces itself.

Generation, Reward, and Optimization (§4.3.1 to §4.3.3)

§4.3.1 Generation categorizes how candidates are produced: Self-Exploratory (SE) sampling directly from the current policy, Refined (R) generation that iteratively improves an initial response, and Interactive (I) generation driven by collaborative, adversarial, or tool-augmented dynamics.

§4.3.2 Reward categorizes how those candidates are scored: Heuristic (H) signals from consistency or hand-designed rules, Model-Based (M) signals from self-evaluation or external judges and reward models, and Verification (V) signals from ground-truth matching, formal provers, or code execution.

§4.3.3 Optimization follows directly from the reward format: SFT on filtered samples, DPO-style objectives on preference orderings, PPO/GRPO on scalar rewards, or specialized self-play objectives.

The table below organizes SGO methods by their generation strategy, reward type, and optimization method (mirroring Table 6 of the paper). Abbreviations: SE = Self-Exploratory, R = Refined, I = Interactive; H = Heuristic, M = Model-Based, V = Verification. & indicates that multiple techniques are used within the same stage.

MethodGenerationRewardOptimization
(2022, Mar) [NeurIPS 2022] STaR: Bootstrapping Reasoning With ReasoningSEVSFT
(2023, May) [ICLR 2024] SIRLC: Language Model Self-improvement by Reinforcement Learning ContemplationSEMPPO
(2023, Oct) SELF: Self-Evolution with Language FeedbackRM & VSFT
(2023, Dec) [EMNLP 2023] LMSI: Large Language Models Can Self-ImproveSEHSFT
(2024, Jan) [ICML 2024] SPIN: Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsSEHSPIN
(2024, May) [ICML 2024] Self-Rewarding Language ModelsSEMDPO
(2024, May) [ICLR 2025] SPPO: Self-Play Preference Optimization for Language Model AlignmentSEMSPPO
(2024, Jun) [NAACL 2024] TRIPOST: Teaching Language Models to Self-Improve through Interactive DemonstrationsRMSFT
(2024, Jun) [NeurIPS 2024] ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchSEMSFT
(2024, Jul) [NeurIPS 2024] RISE: Recursive Introspection: Teaching Language Model Agents How to Self-ImproveRM & VSFT
(2024, Jul) [COLM 2024] V-STaR: Training Verifiers for Self-Taught ReasonersSEVSFT & DPO
(2024, Jul) Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeSEMDPO
(2024, Aug) [AAAI 2025] IWSI: Importance Weighting Can Help Large Language Models Self-ImproveSEH & MSFT
(2024, Sep) [ICLR 2025] SCoRe: Training Language Models to Self-Correct via Reinforcement LearningRVMulti-turn RL
(2024, Oct) [ICLR 2025] ReGenesis: LLMs can Grow into Reasoning Generalists via Self-ImprovementSE & RH & VSFT
(2024, Oct) [ICLR 2025] SynPO: Self-Boosting Large Language Models with Synthetic Preference DataRHDPO
(2024, Nov) [ICML 2025] ScPO: Self-Consistency Preference OptimizationSEHScPO
(2025, Jan) [ICLR 2025] Multiagent Finetuning: Self Improvement with Diverse Reasoning ChainsIHSFT
(2025, Feb) [NeurIPS 2025] SiriuS: Self-improving Multi-agent Systems via Bootstrapped ReasoningIVSFT
(2025, Feb) DNPO: Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic DataSEMDNPO
(2025, Feb) [COLM 2025] Goedel-Prover: A Frontier Model for Open-Source Automated Theorem ProvingSEVSFT & DPO
(2025, Feb) SOL-VER: Learning to Solve and Verify: A Self-Play Framework for Code and Test GenerationSEM & VSFT & DPO
(2025, Feb) RSPO: Regularized Self-Play Alignment of Large Language ModelsSEMRSPO
(2025, Feb) Self-Rewarding Correction for Mathematical ReasoningRVSFT & RL
(2025, Mar) STaSC: Self-Taught Self-Correction for Small Language ModelsRVSFT
(2025, Mar) SPHERE: Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language ModelsSE & RM & VDPO
(2025, Apr) [NAACL 2025] SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized DomainsSEHSFT
(2025, May) Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement LearningRVGRPO
(2025, May) RLSR: Reinforcement Learning from Self RewardSEMGRPO
(2025, May) [EMNLP 2025] DTE: Debate, Train, Evolve: Self Evolution of Language Model ReasoningIHSFT & GRPO
(2025, May) [NeurIPS 2025] TBV: Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsSEVPPO
(2025, May) [NeurIPS 2025] SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited DataSEVGRPO
(2025, May) [NeurIPS 2025] Absolute Zero: Reinforced Self-play Reasoning with Zero DataIH & VTRR++
(2025, May) [NeurIPS 2025] STaPLe: Latent Principle Discovery for Language Model Self-ImprovementRHSFT
(2025, Jun) [NeurIPS 2025] Self-Challenging Language Model AgentsIVSFT
(2025, Jun) PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative VerifierRVMulti-turn RL
(2025, Jun) [NeurIPS 2025] CURE: Co-Evolving LLM Coder and Unit Tester via Reinforcement LearningIH & VPPO
(2025, Jul) [ACL 2025 Findings] CRESCENT: The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?SEHSFT
(2025, Jul) [ACL 2025 Findings] Unlocking LLMs' Self-Improvement Capacity with Autonomous Learning for Domain AdaptationRHDPO
(2025, Aug) [ICLR 2026] R-Zero: Self-Evolving Reasoning LLM from Zero DataIH & VGRPO
(2025, Sep) Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended TasksSEH & MSFT
(2025, Oct) [ICLR 2026] RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-PenalizationSEHGRPO
(2025, Oct) SPICE: Self-Play In Corpus Environments Improves ReasoningIH & VDrGRPO
(2025, Oct) MAE: Multi-Agent Evolve: LLM Self-Improve through Co-evolutionIMGRPO
(2025, Dec) SSR: Toward Training Superintelligent Software Agents through Self-Play SWE-RLIH & VSWE-RL
(2026, Jan) [ICLR 2026] ReVeal: Self-Evolving Code Agents via Reliable Self-VerificationIH & VTAPO
(2026, Jan) Dr. Zero: Self-Evolving Search Agents without Training DataIVGRPO & HRPO
(2026, Jul) [ACL 2026 Findings] EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning for LLMsSEVGRPO

Representative Instances of SGO (§4.3.4)

§4.3.4 Representative Instances of SGO highlights the few recurring patterns through which most methods instantiate the loop, defining the structural relationship between generation, reward, and optimization.

Iterative Rejection Sampling. The model generates diverse candidates, filters them via ground truth (oracle) or statistical consistency (majority vote), and fine-tunes on the retained pseudo-labels. Improvement comes from distilling the model's own best generations into its weights.

Self-Verification & Refinement. The model actively evaluates, scores, or refines its own outputs using self-generated reward signals. Unlike rejection sampling, the model plays a semantic role as its own judge, and updates via RL or DPO.

Self-Play. The model improves through dynamic interaction between multiple roles (adversarial proposer-solver or collaborative debate), providing an evolving curriculum of challenges that pushes beyond its initial distribution.

Self-Evolving Optimization (§4.4)

§4.4 Self-Evolving Optimization moves the target of improvement from the model's parameters to the optimization process itself: the update rule, the learning algorithm, or the surrounding agentic scaffolding is revised, so that the way the model improves can itself improve. This extends the unit of improvement from a single model's weights to the whole agentic system.

Theoretical Analysis (§4.5.1)

§4.5.1 Theoretical Analysis probes the premise that a model can become better by training on data it generated itself, along three themes: where the loop's improvement originates (the "sharpening" mechanism and the generation-verification gap), whether it converges and to what ceiling, and when it collapses.

Inference Refinement (§5)

§5 Inference Refinement focuses on improving the model's output quality during inference without permanently updating its parameters. It spans decoding-level enhancements, structured reasoning processes, agentic system coordination, and test-time parameter adaptation.

Decoding Strategies (§5.2)

§5.2 Decoding Strategies explicitly guide output generation at the token or sequence level to steer the model toward higher-quality outputs by modifying how candidates are sampled, searched, scored, or blended.

Sampling-Based (§5.2.1)

§5.2.1 Sampling-Based methods generate multiple candidate outputs and select or synthesize the best one using consistency voting, reward-based ranking, or ensemble fusion.

Tree Search (§5.2.2)

§5.2.2 Tree Search methods organize the decoding process as a tree of partial solutions, using beam search, Monte Carlo Tree Search, or graph-based exploration to systematically evaluate branching reasoning paths.

Logit and Probability Adjustments (§5.2.3)

§5.2.3 Logit and Probability Adjustments modify the model's output distribution at the logit level — through contrastive decoding, reward-guided reranking, or logit blending — to steer generation toward desired properties without retraining.

Efficiency-Oriented Methods (§5.2.4)

§5.2.4 Efficiency-Oriented Methods accelerate inference through speculative decoding, parallel generation, and other techniques that reduce latency while maintaining output quality.

Reasoning-Based Improvement (§5.3)

§5.3 Reasoning-Based Improvement structures the model's intermediate thought process through feedback loops, planning decomposition, and collaborative debate, enabling dynamic multi-step refinement of reasoning at inference time.

Feedback-Based Reasoning (§5.3.1)

§5.3.1 Feedback-Based Reasoning transforms static inference into a dynamic closed-loop process where the model iteratively critiques and refines its outputs using either self-generated evaluations or external verification signals (e.g., code execution, tool outputs).

Planning-Based Reasoning (§5.3.2)

§5.3.2 Planning-Based Reasoning enables the model to decompose complex objectives into smaller sub-goals, either through open-loop planning (a complete blueprint before execution) or closed-loop planning (dynamically adjusting the plan based on environmental feedback).

Collaborative Reasoning (§5.3.3)

§5.3.3 Collaborative Reasoning distributes reasoning across an ensemble of interacting agents — through role specialization, structured debate, or cooperative refinement — so that agents iteratively build upon and correct each other's reasoning traces.

Agentic System Improvement (§5.4)

§5.4 Agentic System-Based Improvement extends inference-time refinement to the system level by dynamically adapting prompts, memory, tool libraries, and workflows — enabling self-improvement not by altering internal generation but by evolving the environment in which the model operates.

Prompts (§5.4.1)

§5.4.1 Prompts covers dynamic prompt optimization at inference time, including sampling-based, search-based, evolutionary, and textual-gradient-descent approaches that refine instructions and in-context demonstrations for better task performance.

Memory (§5.4.2)

§5.4.2 Memory covers persistent long-term and working memory structures that allow agents to accumulate, retrieve, and consolidate knowledge across interactions, enabling adaptive behavior and continual improvement beyond a single context window.

Tooling (§5.4.3)

§5.4.3 Tooling covers how agents dynamically create, refine, select, and manage external tools (APIs, code, documentation) at inference time, transforming tool use from static invocation into a self-expanding procedural ecosystem.

Workflow and System Evolution (§5.4.4)

§5.4.4 Workflow and System Evolution covers how multi-agent systems dynamically evolve their communication topology, coordination protocols, and overall architecture — either before or during deployment — to adapt to task-specific requirements.

Test-Time Training (§5.5)

§5.5 Test-Time Training (TTT) represents a paradigm shift from static inference to dynamic, gradient-based self-improvement at inference time. The model temporarily adapts its parameters to instance-specific challenges on the fly, blurring the boundary between training and deployment.

TT-SFT (§5.5)

TT-SFT (Test-Time Supervised Fine-Tuning) updates model parameters during inference using a supervised loss derived from instance-specific data, encoding structural patterns from retrieved examples or self-generated candidates directly into the model's weights.

TT-RL (§5.5)

TT-RL (Test-Time Reinforcement Learning) performs full RL updates on unlabeled data at test time, using majority vote or reward model signals as self-supervision to adapt the model's policy on the fly.

Autonomous Evaluation (§6)

§6 Autonomous Evaluation provides continuous feedback on model performance to steer the self-improvement cycle. It moves beyond static benchmarks to dynamic evaluation sets and interactive environments that can keep pace with evolving model capabilities.

Dynamic Benchmarking (§6.2)

§6.2 Dynamic Benchmarking continuously regenerates or transforms evaluation instances to mitigate data contamination and distributional staleness, ensuring that benchmarks remain informative as models improve through iterative self-training.

Interactive Environment Evaluation (§6.3)

§6.3 Interactive Environment Evaluation embeds the model within interactive, stateful environments where performance is assessed over execution trajectories rather than isolated input-output pairs, directly probing planning, recovery, and long-horizon behavioral consistency. It includes outcome-based (§6.3.1) environments that evaluate terminal goal satisfaction, and process-based (§6.3.2) environments that additionally evaluate execution quality.

Challenges and Limitations (§7)

§7 Challenges and Limitations systematically examines the failure modes and bottlenecks that threaten the stability and effectiveness of self-improvement systems, from data degradation to flawed feedback, optimization pathologies, and the limits of autonomous evaluation and supervision.

Data Autophagy (§7.1)

§7.1 Data Autophagy describes the degradation of information quality within self-improvement loops: as systems reuse their own synthetic outputs, errors accumulate through data copying, catastrophic forgetting, and model collapse, leading to progressive loss of diversity and capability.

Data Collection Limits

Data Copying

Catastrophic Forgetting

Model Collapse

Flawed Feedback Signals (§7.2)

§7.2 Flawed Feedback Signals examines the inherent quality issues in self-generated feedback: bias amplification through iterative loops, inconsistent and unstable evaluation signals, and how these signal-level defects undermine both training-time and inference-time improvement.

Bias Amplification

Feedback Inconsistency and Instability

Optimization-Driven Failures (§7.3)

§7.3 Optimization-Driven Failures examines how strong optimization pressure on imperfect proxy rewards leads to Goodhart's Law scenarios: reward hacking where models exploit reward function loopholes, and emergent misalignment where narrow task optimization generalizes into broader dangerous behaviors.

Reward Hacking

Emergent Misalignment

Ineffective Self-Refinement (§7.4)

§7.4 Ineffective Self-Refinement examines how inference-time feedback loops fail in practice: the generation-verification gap may be illusory (models are not reliably better at judging than generating), and self-correction often degrades rather than improves outputs.

Illusory Generation-Verification Gap

Limits of Self-Correction

Evaluation Bottlenecks (§7.5)

§7.5 Evaluation Bottlenecks questions whether current benchmarks, metrics, and LLM-based evaluators can reliably measure self-improvement, covering benchmark contamination, metric design deficiencies, and systematic biases in LLM-as-a-judge systems.

Benchmark Contamination

Metric Design Deficiencies

LLM Evaluator Bias

Supervision Bottlenecks (§7.6)

§7.6 Supervision Bottlenecks reveals the limits of maintaining external control over self-improvement systems: human supervision quality degrades as models scale, and even accurate supervision signals may be ineffective due to alignment faking and controllability limitations.

Limited Supervision Quality

Supervision Ineffectiveness

Budget Constraints (§7.7)

§7.7 Budget Constraints examines the practical feasibility of self-improvement under finite compute. Costs compound from two sources: iterated compute costs, since generation, selection, reward computation, optimization, and evaluation are repeated in full at every round; and strong-model dependence, since the quality of self-generated data and self-assigned rewards is bounded by the base model's own capability.

Iterated Compute Costs

Strong-Model Dependence

Potential Risks (§8)

§8 Potential Risks examines the safety concerns that arise even when a self-improvement system works exactly as intended. Whereas §7 covers failure modes at individual stages, this section asks what a well-functioning autonomous loop can put at risk, along four dimensions: loss of control, high-stakes harm, social and ethical risks, and misuse risks.

Loss of Control (§8.1)

§8.1 Loss of Control arises when autonomous and recursive self-improvement lets a model raise the very capabilities that drive its own optimization, so human control over the system progressively weakens. It covers oversight loss (supervisors can no longer detect, constrain, or reverse harmful behavior) and deceptive alignment (the model behaves as intended only while it judges it is being watched).

Oversight Loss

Deceptive Alignment

High-Stakes Harm (§8.2)

§8.2 High-Stakes Harm arises when a self-improving agent is deployed in a domain that tolerates little error, such as clinical care or financial markets. A loop that keeps rewriting its own behavior can compound small errors into concrete harm before oversight intervenes, producing individual harm (one autonomous decision injures the person it affects) and systemic failure (coupled agents destabilize the system they operate in).

Individual Harm

Systemic Failure

Social and Ethical Risks (§8.3)

§8.3 Social and Ethical Risks arise when self-improvement systems are deployed at scale and outpace the institutions, norms, and accountability structures meant to govern them. It covers social disruption (labor displacement, power concentration, degradation of the information commons) and ethical breakdown (the responsibility gap, value lock-in, and erosion of human agency).

Social Disruption

Ethical Breakdown

Misuse Risks (§8.4)

§8.4 Misuse Risks concern the ways a self-improving system can harm the user and the systems it acts on, or be turned into an instrument of attack, because it couples autonomous optimization with broad tool, code, and data access. It covers unintended damage from ordinary goal pursuit and deliberate misuse by adversarial operators or by the model optimizing its own harmful capability.

Unintended Damage

Deliberate Misuse

Applications (§9)

§9 Applications surveys how self-improvement mechanisms are applied across six domains, enabling specialized models and self-evolving agents to iteratively refine their expertise within constrained environments.

Code

Coding agents achieve self-evolution by exploiting the definitive feedback from compilers and unit tests across the software development lifecycle.

Math

Mathematical self-evolution relies on formal logical consistency and rigorous verification of reasoning paths, from code-assisted execution to formal theorem proving.

Healthcare

Self-evolving systems in healthcare emphasize clinical safety, diagnostic accuracy, and multidisciplinary collaboration.

Finance

Financial agents evolve their strategies in high-noise environments by utilizing layered memory and risk-sensitive simulation.

Science

Scientific applications center on self-improving agents that autonomously make discoveries, spanning natural-science findings, novel algorithms, and improved agent designs.

Auto Research

Whereas the systems above target discovery, a distinct line pursues end-to-end research automation, where a single agentic system carries a project through the full scholarly lifecycle: ideation, experimental design and execution, manuscript writing, and validation.


Future Outlook (§10)

1. From Model-Level Optimization to End-to-End Self-Improving Systems

Future work may move beyond improving individual components toward end-to-end self-improvement system that continuously generate data, evaluate outputs, and update themselves within an automated loop.

2. Toward Specialized and Application-Centric Self-Improving Models

Self-improvement mechanisms are increasingly applied in domain-specific settings such as coding, science, finance, and healthcare, enabling specialized agents to iteratively refine expertise within constrained environments.

3. Unified Benchmarks for Self-Improvement and Autonomous Evaluation

The field still lacks standardized benchmarks designed to measure iterative improvement, stability across cycles, and long-term capability growth.

4. Balancing Automation and Human Oversight

Future systems must balance autonomous improvement with human supervision to ensure scalability while maintaining safety, reliability, and alignment. The risks analyzed in §8, from oversight loss and deceptive alignment to high-stakes harm and misuse, arise precisely when human control over the improvement loop weakens.

Citation

If you find this work useful, please cite:

@misc{yang2026selfimprovementlargelanguagemodels,
      title={Self-Improvement of Large Language Models: A Technical Overview and Future Outlook},
      author={Haoyan Yang and Mario Xerri and Solha Park and Huajian Zhang and Yiyang Feng and Sai Akhil Kogilathota and Jiawei Zhou},
      year={2026},
      eprint={2603.25681},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2603.25681}
}

Team

This survey is authored by members of the Zesearch NLP Lab at Stony Brook University:

We are interested in the self-improvement of LLMs and recursive self-improvement (RSI). If you have any questions or ideas, feel free to reach out.