Repository for Self-Evolving Coding Agents
See the code
Overview of self-evolving coding agents.
Coding agents increasingly learn from execution outcomes, trajectories, accumulated experience, and environmental feedback, improving the persistent components of their own software-engineering workflow. This repository accompanies our survey and curates its paper corpus, related methods, benchmarks, products, and related surveys. Self-evolving systems are organized into three layers: assets (memory, skills, tools, context), architecture (harness, workflow, multi-agent structures), and model weights.
This repository covers five groups of resources:
Self-Evolving Coding Agents: curated systems organized by which part of the agent gets modified: assets (memory, skills, tools, context), architecture (harness, workflow, topology), or model weights.
General Self-Evolution Methods in Coding Settings: General agent self-evolution methods whose improvements are evaluated on code generation, program execution, or software engineering tasks.
Benchmarks and Empirical Studies: Conventional benchmarks for repository-level software engineering and general coding, dedicated benchmarks for agent self-evolution, and empirical studies of coding-agent self-evolution.
Products: deployed coding products with persistent adaptation mechanisms, mapped to the same target vocabulary.
Related Surveys: surveys covering self-evolving agents, coding agents, and their intersection.
[2025-arXiv][2025-arXiv] · [Code][2024-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-FSE] · [Code][2026-arXiv][2026-arXiv][2026-ICLR] · [Code][2023-NeurIPS] · [Code][2025-NeurIPS] · [Code][2025-ICML Workshop] · [Code][2025-ICML] · [Code][2026-arXiv][2026-arXiv][2025-REALM] · [Code][2026-ICLR][2026-arXiv][2026-arXiv][2026-IEEE ICE][2026-arXiv][2026-ACM CAIS][2026-arXiv][2026-arXiv][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2026-arXiv][2026-ASE] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv][2025-arXiv] · [Code][2026-arXiv][2024-EMNLP Findings] · [Code][2026-ACL Findings] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-TOSEM] · [Code][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2026-ICLR] · [Code][2026-OpenReview] · [Code][2026-ICLR] · [Code][2026-ICLR][2026-arXiv][2026-arXiv] · [Code][2025-ICLR] · [Code][2024-COLM] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-ACL] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2025-ICLR] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2025-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2025-ICLR] · [Code][2025-arXiv] · [Code][2025-arXiv][2025-EMNLP Demos] · [Code][2026-arXiv] · [Code][2026-arXiv][2025-ICLR] · [Code][2026-ICML][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2026-ICML][2026-ICLR] · [Code][2025-NeurIPS] · [Code][2026-arXiv][2025-NeurIPS][2026-arXiv][2026-ICML][2025-NeurIPS] · [Code][2026-arXiv] · [Code][2026-ICLR][2025-NeurIPS] · [Code][2025-ICML] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code]This section brings together repository-level software engineering benchmarks and general coding benchmarks to assess coding capabilities, self-evolution benchmarks to evaluate agents' ability to improve themselves, and empirical studies to examine the performance gains, computational costs, and failure modes of self-evolving coding agents.
[2024-ICLR][2026-ICML][2025-arXiv][2025-ICLR][2025-NeurIPS Datasets and Benchmarks][2025-arXiv][2025-NeurIPS Datasets and Benchmarks][2025-Benchmark][2026-arXiv][2025-arXiv][2024-ICSE][2024-ACL Findings][2026-ACL Findings][2025-COLM][2025-COLM][2025-arXiv][2025-NeurIPS][2022-MSR][2025-Dataset][2025-ICLR][2026-arXiv][2026-arXiv][2026-Benchmark][2025-Benchmark][2021-arXiv][2021-arXiv][2021-NeurIPS Datasets and Benchmarks][2022-Science][2025-ICLR][2025-ICLR][2023-NeurIPS][2023-IEEE TSE][2023-ICML][2024-ICML][2024-ACL][2025-NeurIPS Datasets and Benchmarks][2024-Benchmark][2026-ICLR][2025-arXiv][2026-arXiv][2024-ACL Findings][2026-ICML][2025-arXiv][2023-ICML][2024-EMNLP Findings][2023-NeurIPS Datasets and Benchmarks][2025-ICLR][2024-EMNLP][2023-NeurIPS Datasets and Benchmarks][2023-ICCAD][2024-ASP-DAC][2025-arXiv][2025-ICML][2026-arXiv][2025-ICLR][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2025-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-ACL]These products and open-source tools support persistent memory, reusable skills, context updates, and harness customization through automatic or user-guided refinement. The table summarizes adaptation targets and mechanisms. Inclusion reflects documented capabilities, not necessarily a fully autonomous, experimentally validated self-evolution loop.
Target legend: 🔵 Assets — Memory, Skill, Tool, Context (including environment adaptation) · 🟣 Architecture — Harness, Workflow, Multi-Agent.
| Product | Company / Year | Target | Mechanism in Practice | Resources |
|---|---|---|---|---|
| Prime Agent (PA) | Prime Intellect2026 | 🔵 AssetsMemory · Skill · Context🟣 Architecture Harness · Multi-Agent | Refines agent components from task trajectories. | Code |
| DeepSeek Harness (DSH) | DeepSeek2026 | 🟣 ArchitectureHarness | Tests plugins and creates presets through Creator Mode. | Code |
| Gemini CLI Auto Memory (GCAM) | Google2026 | 🔵 AssetsMemory · Skill | Proposes memory and skill updates for user review. | Changelog |
| GitHub Copilot Memory (GCM) | GitHub2026 | 🔵 AssetsMemory | Stores repository facts and validates them before reuse. | Announcement |
| Augment Agent / Cosmos Learning Flywheel (AA/CLF) | Augment Code2025–2026 | 🔵 AssetsMemory · Context | Distills team feedback into shared memory and context. | Memory review |
| Claude Code Auto Memory (CCAM) | Anthropic2026 | 🔵 AssetsMemory | Automatically saves corrections, preferences, and project learnings. | Changelog |
| Cursor Memories / Automations (CMA) | Cursor2025–2026 | 🔵 AssetsMemory | Retains useful context across sessions and automation runs. | Memories |
| Devin Session Insights / Knowledge / Playbooks (Devin SIKP) | Cognition2025–2026 | 🔵 AssetsMemory · Skill · Context | Uses session analysis to guide knowledge and configuration updates. | Advanced capabilities |
| Windsurf Cascade Memories (WCM) | Windsurf / Cognition2025–2026 | 🔵 AssetsMemory | Creates workspace memories and retrieves them when relevant. | Docs |
| OpenBlock Agent (OB-1) | OpenBlock Labs2026 | 🔵 AssetsMemory | Learns codebase patterns for subsequent tasks. | Waitlist |
| Letta Code | Letta2025–2026 | 🔵 AssetsMemory · Skill · Context | Reviews sessions to refine memory and context; versions skills. | Code |
| Hermes Agent | Nous Research2026 | 🔵 AssetsMemory · Skill | Creates skills from experience and refines them through reuse. | Skills Code |
| Kiro Web / Autonomous Agent | AWS2025–2026 | 🔵 AssetsMemory | Learns team conventions from code-review feedback. | Product |
| Replit Agent | Replit2025–2026 | 🔵 AssetsMemory · Skill · Context | Maintains replit.md and creates reusable skills. | Announcement |
| OpenClaw | OpenClaw Foundation / Community2026 | 🔵 AssetsMemory | Consolidates session notes into persistent memory. | Coding integration Code |
| Cline Memory Bank | Cline / Community2025 | 🔵 AssetsMemory · Context | Maintains structured project memory through configured instructions. | Documentation |
From memory to architecture: Most products adapt assets such as memory, skills, and context. Prime Agent and DeepSeek Harness also expose mechanisms for modifying the surrounding agent harness.
[2026-TMLR][2025-arXiv][2026-TechRxiv][2026-arXiv][2026-OpenReview][2024-arXiv][2026-arXiv][2026-OpenReview][2026-Preprints.org][2026-arXiv][2025-TOSEM][2025-Automated Software Engineering][2025-TOSEM][2024-TOSEM][2026-TOSEM][2026-arXiv]@misc{zhou2026selfevolvingcodingagents,
title={Self-Evolving Coding Agents},
author={Hao Zhou and Haichuan Hu and Tianyu Luo and Ye Shang and Chunrong Fang and Zhenyu Chen and Liang Xiao and Quanjun Zhang},
year={2026},
eprint={2608.03392},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2608.03392},
}
🤝 Contributions are welcome! If you find any missing or incorrect information, please feel free to open an issue or submit a pull request.
Repository for Self-Evolving Coding Agents
See the code
Overview of self-evolving coding agents.
Coding agents increasingly learn from execution outcomes, trajectories, accumulated experience, and environmental feedback, improving the persistent components of their own software-engineering workflow. This repository accompanies our survey and curates its paper corpus, related methods, benchmarks, products, and related surveys. Self-evolving systems are organized into three layers: assets (memory, skills, tools, context), architecture (harness, workflow, multi-agent structures), and model weights.
This repository covers five groups of resources:
Self-Evolving Coding Agents: curated systems organized by which part of the agent gets modified: assets (memory, skills, tools, context), architecture (harness, workflow, topology), or model weights.
General Self-Evolution Methods in Coding Settings: General agent self-evolution methods whose improvements are evaluated on code generation, program execution, or software engineering tasks.
Benchmarks and Empirical Studies: Conventional benchmarks for repository-level software engineering and general coding, dedicated benchmarks for agent self-evolution, and empirical studies of coding-agent self-evolution.
Products: deployed coding products with persistent adaptation mechanisms, mapped to the same target vocabulary.
Related Surveys: surveys covering self-evolving agents, coding agents, and their intersection.
[2025-arXiv][2025-arXiv] · [Code][2024-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-FSE] · [Code][2026-arXiv][2026-arXiv][2026-ICLR] · [Code][2023-NeurIPS] · [Code][2025-NeurIPS] · [Code][2025-ICML Workshop] · [Code][2025-ICML] · [Code][2026-arXiv][2026-arXiv][2025-REALM] · [Code][2026-ICLR][2026-arXiv][2026-arXiv][2026-IEEE ICE][2026-arXiv][2026-ACM CAIS][2026-arXiv][2026-arXiv][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2026-arXiv][2026-ASE] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv][2025-arXiv] · [Code][2026-arXiv][2024-EMNLP Findings] · [Code][2026-ACL Findings] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-TOSEM] · [Code][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2026-ICLR] · [Code][2026-OpenReview] · [Code][2026-ICLR] · [Code][2026-ICLR][2026-arXiv][2026-arXiv] · [Code][2025-ICLR] · [Code][2024-COLM] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-ACL] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2025-ICLR] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2025-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code][2025-ICLR] · [Code][2025-arXiv] · [Code][2025-arXiv][2025-EMNLP Demos] · [Code][2026-arXiv] · [Code][2026-arXiv][2025-ICLR] · [Code][2026-ICML][2026-arXiv][2026-arXiv] · [Code][2026-arXiv][2026-arXiv] · [Code][2026-ICML][2026-ICLR] · [Code][2025-NeurIPS] · [Code][2026-arXiv][2025-NeurIPS][2026-arXiv][2026-ICML][2025-NeurIPS] · [Code][2026-arXiv] · [Code][2026-ICLR][2025-NeurIPS] · [Code][2025-ICML] · [Code][2026-arXiv] · [Code][2026-arXiv] · [Code]This section brings together repository-level software engineering benchmarks and general coding benchmarks to assess coding capabilities, self-evolution benchmarks to evaluate agents' ability to improve themselves, and empirical studies to examine the performance gains, computational costs, and failure modes of self-evolving coding agents.
[2024-ICLR][2026-ICML][2025-arXiv][2025-ICLR][2025-NeurIPS Datasets and Benchmarks][2025-arXiv][2025-NeurIPS Datasets and Benchmarks][2025-Benchmark][2026-arXiv][2025-arXiv][2024-ICSE][2024-ACL Findings][2026-ACL Findings][2025-COLM][2025-COLM][2025-arXiv][2025-NeurIPS][2022-MSR][2025-Dataset][2025-ICLR][2026-arXiv][2026-arXiv][2026-Benchmark][2025-Benchmark][2021-arXiv][2021-arXiv][2021-NeurIPS Datasets and Benchmarks][2022-Science][2025-ICLR][2025-ICLR][2023-NeurIPS][2023-IEEE TSE][2023-ICML][2024-ICML][2024-ACL][2025-NeurIPS Datasets and Benchmarks][2024-Benchmark][2026-ICLR][2025-arXiv][2026-arXiv][2024-ACL Findings][2026-ICML][2025-arXiv][2023-ICML][2024-EMNLP Findings][2023-NeurIPS Datasets and Benchmarks][2025-ICLR][2024-EMNLP][2023-NeurIPS Datasets and Benchmarks][2023-ICCAD][2024-ASP-DAC][2025-arXiv][2025-ICML][2026-arXiv][2025-ICLR][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2025-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-arXiv][2026-ACL]These products and open-source tools support persistent memory, reusable skills, context updates, and harness customization through automatic or user-guided refinement. The table summarizes adaptation targets and mechanisms. Inclusion reflects documented capabilities, not necessarily a fully autonomous, experimentally validated self-evolution loop.
Target legend: 🔵 Assets — Memory, Skill, Tool, Context (including environment adaptation) · 🟣 Architecture — Harness, Workflow, Multi-Agent.
| Product | Company / Year | Target | Mechanism in Practice | Resources |
|---|---|---|---|---|
| Prime Agent (PA) | Prime Intellect2026 | 🔵 AssetsMemory · Skill · Context🟣 Architecture Harness · Multi-Agent | Refines agent components from task trajectories. | Code |
| DeepSeek Harness (DSH) | DeepSeek2026 | 🟣 ArchitectureHarness | Tests plugins and creates presets through Creator Mode. | Code |
| Gemini CLI Auto Memory (GCAM) | Google2026 | 🔵 AssetsMemory · Skill | Proposes memory and skill updates for user review. | Changelog |
| GitHub Copilot Memory (GCM) | GitHub2026 | 🔵 AssetsMemory | Stores repository facts and validates them before reuse. | Announcement |
| Augment Agent / Cosmos Learning Flywheel (AA/CLF) | Augment Code2025–2026 | 🔵 AssetsMemory · Context | Distills team feedback into shared memory and context. | Memory review |
| Claude Code Auto Memory (CCAM) | Anthropic2026 | 🔵 AssetsMemory | Automatically saves corrections, preferences, and project learnings. | Changelog |
| Cursor Memories / Automations (CMA) | Cursor2025–2026 | 🔵 AssetsMemory | Retains useful context across sessions and automation runs. | Memories |
| Devin Session Insights / Knowledge / Playbooks (Devin SIKP) | Cognition2025–2026 | 🔵 AssetsMemory · Skill · Context | Uses session analysis to guide knowledge and configuration updates. | Advanced capabilities |
| Windsurf Cascade Memories (WCM) | Windsurf / Cognition2025–2026 | 🔵 AssetsMemory | Creates workspace memories and retrieves them when relevant. | Docs |
| OpenBlock Agent (OB-1) | OpenBlock Labs2026 | 🔵 AssetsMemory | Learns codebase patterns for subsequent tasks. | Waitlist |
| Letta Code | Letta2025–2026 | 🔵 AssetsMemory · Skill · Context | Reviews sessions to refine memory and context; versions skills. | Code |
| Hermes Agent | Nous Research2026 | 🔵 AssetsMemory · Skill | Creates skills from experience and refines them through reuse. | Skills Code |
| Kiro Web / Autonomous Agent | AWS2025–2026 | 🔵 AssetsMemory | Learns team conventions from code-review feedback. | Product |
| Replit Agent | Replit2025–2026 | 🔵 AssetsMemory · Skill · Context | Maintains replit.md and creates reusable skills. | Announcement |
| OpenClaw | OpenClaw Foundation / Community2026 | 🔵 AssetsMemory | Consolidates session notes into persistent memory. | Coding integration Code |
| Cline Memory Bank | Cline / Community2025 | 🔵 AssetsMemory · Context | Maintains structured project memory through configured instructions. | Documentation |
From memory to architecture: Most products adapt assets such as memory, skills, and context. Prime Agent and DeepSeek Harness also expose mechanisms for modifying the surrounding agent harness.
[2026-TMLR][2025-arXiv][2026-TechRxiv][2026-arXiv][2026-OpenReview][2024-arXiv][2026-arXiv][2026-OpenReview][2026-Preprints.org][2026-arXiv][2025-TOSEM][2025-Automated Software Engineering][2025-TOSEM][2024-TOSEM][2026-TOSEM][2026-arXiv]@misc{zhou2026selfevolvingcodingagents,
title={Self-Evolving Coding Agents},
author={Hao Zhou and Haichuan Hu and Tianyu Luo and Ye Shang and Chunrong Fang and Zhenyu Chen and Liang Xiao and Quanjun Zhang},
year={2026},
eprint={2608.03392},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2608.03392},
}
🤝 Contributions are welcome! If you find any missing or incorrect information, please feel free to open an issue or submit a pull request.