Awesome Agentic Evolution

From self-improving agents to co-evolving, open-ended agent ecosystems.
A community-curated index of agents that preserve improvements across attempts,
sessions, or generations. The collection maps what changes, what feedback drives
the change, what persists, and how the claimed improvement is evaluated.
Explore the living research dashboard →
Last editorial review: 2026-09-20 (structure and selected sources; not a full evidence audit)
Writing a survey? Begin with the survey workspace,
the complete inventory and extraction status, and the
classification rules. The repository is an evolving evidence
base, not yet a completed systematic review or a verified living survey.
Contents
Scope
Agentic evolution requires a persistent update: an artifact changed during one
improvement cycle must influence future behavior. A resource is in scope when
it makes at least three elements legible:
- the artifact being updated;
- the feedback, reward, evaluator, or environmental signal driving the update;
- the retained result and an evaluation showing whether it helped.
Generic agent frameworks, static retrieval systems, and one-shot
self-correction are out of scope unless they participate in a persistent
improvement loop. Benchmarks and evaluators are listed separately because they
measure evolution but do not necessarily evolve themselves.
Taxonomy
The primary classification question is which persistent artifact is edited,
not which mechanism a paper uses or which application it studies.
- Parameters: model weights, policies, adapters, or other trainable state.
- Memory: agent-specific traces, reflections, summaries, and experience
records retained across future attempts.
- Knowledge: external or decontextualized world knowledge, retrieval
corpora, knowledge bases, schemas, and reusable factual assets.
- Skills: reusable procedural instructions, routines, and callable
strategies for completing classes of tasks.
- Tools: executable interfaces, APIs, code modules, and expert agents the
agent can invoke.
- Topology: prompts, routing, control flow, graph connectivity, role
decomposition, and run-time orchestration.
- Co-evolution: coupled change in the agent and one or more external factors:
objectives, evaluators, environments, update mechanisms, or populations.
Resources may have multiple targets, but each is listed once under its primary
target. The Targets line records other independently editable artifacts.
Implementation code is not a separate category: classify the behavioral
artifact that the code changes—for example, a generated tool under Tools or
a rewritten control graph under Topology.
Boundary rule: self-history belongs to Memory; externally reusable world
information belongs to Knowledge. When a system bundles both, annotate both
only if they can be updated independently.
Co-evolution is a cross-cutting relation, not a seventh internal artifact.
Use its section when coupled adaptation is the central contribution; put
Co-evolution first there and list the internal targets afterwards. Merely
using tools, memory, or multiple agents does not mean those objects evolve.
See definitions, counterexamples, and classification decisions.
Start Here
Read these surveys for orientation, then compare their coverage using the
survey questions and outline.
They are background sources, not additional method entries.
Survey Workspace
The inventory is generated from this README. Detailed reviews live separately;
missing reviews remain explicitly unextracted. Link checks, source reading,
full evidence review, and experimental reproduction are different activities.
Resource Map
Entries appear once under their primary persistent target and are alphabetized
within each category. Paper, code, and article links for the same work stay
together. These are indexed resources, not a count of survey-verified papers.
Repository-only implementations and paper-linked work are distinguished in the
structured inventory; results are author-reported
unless an evidence record explicitly documents independent reproduction.
Parameters
- AgentEvolver — Couples task
generation, experience synthesis, and reinforcement learning for persistent
agent improvement.
Targets: Parameters, Memory, Co-evolution.
- ARISE-RL: Reward-Gated Self-Evolution —
Co-evolves a rubric/task generator and tool-using solver, gating
memory-augmented self-distillation on reward improvement across
deep-research and travel-planning tasks.
Targets: Parameters, Co-evolution.
- Astar — Trains an evolution-guiding
model from industrial iteration histories, uses a surrogate reward
evaluator, and guides 20 consecutive iterations with offline and online
gains.
Targets: Parameters.
- Benchmark-as-Teacher
— Paper. Uses held-out-safe stage states
and stage rubrics to select curricula, updates the policy with GRPO, and
gates each checkpoint on the fixed benchmark contract.
Targets: Parameters, Co-evolution.
- CAFE: Self-Improving Search Agents Need Co-Evolving Feedback — Co-evolves a search agent and
critic from matched failures, alternating online and offline feedback
updates with gains across seven search benchmarks and six out-of-domain
sets.
Targets: Parameters, Co-evolution.
- FlowBalance — Reweights policy updates
from verifier-derived group advantage and calibrated self-guidance,
retaining or reversing on-policy signals to improve reasoning reward and
training stability.
Targets: Parameters.
- GenRubric —
Paper. Uses rubric-induced
self-consistency rewards to evolve rubric-generation parameters from
unlabeled queries, with agreement gains on human-annotated benchmarks and
held-out domains.
Targets: Parameters.
- RecurSE: Bounded Recursive Self-Evaluation for LLM Rubric Judges — Co-evolves a rubric judge and
checker with decoupled self-reward, validation-based early stopping,
held-out generalization, and ablations against frozen-checker and teacher
baselines.
Targets: Parameters, Co-evolution.
- Self-Rewarding Language Models — Uses
the model as both instruction follower and judge during iterative training.
Targets: Parameters.
- STaR — Iteratively generates successful
rationales and updates model parameters on the retained examples.
Targets: Parameters.
- WebWorld — Uses browser-issued
acceptance certificates to ratchet verified web-code transitions into
training data, improving HTMLBench and MiniAppBench under matched training.
Targets: Parameters.
Memory
- A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks — Abstracts successful jailbreaks
into reusable method-level rules, retains them across interactions, and
reduces attack success across four families while preserving benign utility.
Targets: Memory.
- APEx: Distillation of Agent Procedural Experience — Distills trajectory memories
into procedural skills, then adapts a research planner with reward-guided
test-time reinforcement learning across seven benchmarks.
Targets: Memory, Parameters.
- Areev — Uses deterministic,
evidence-citing analyzers to propose memory changes, requires named
approval, stores inverse records, and remeasures outcomes so regressions can
be reverted.
Targets: Memory, Tools.
- AutoMem — Searches task-adaptive memory
architectures from historical trajectories and failure-guided module
feedback, reporting gains across three benchmarks and two backbones.
Targets: Memory, Topology.
- CHIME: Credit-Aware Hierarchical Memory Evolution — Separates planning and
execution memory banks, attributes outcomes before retention, and transfers
credit-aware memories across four long-horizon benchmarks and backbone
models.
Targets: Memory.
- CrystalMem — Reversibly demotes and
verifies memory entries under changing byte budgets, using influence-aware
retention and evaluation across seven environments, seventeen methods, and
six backbones.
Targets: Memory.
- DiagEvo — Extracts recurring failure
causes into hierarchical error memory, targets challenger generation at
active weaknesses, and improves nine benchmarks across three solvers.
Targets: Memory.
- earcon — A local
OpenAI-compatible proxy that distills session outcomes into SQLite
experience cards and injects them into future tasks, with a discrete-action
maze evaluation.
Targets: Memory.
- Evolve —
Paper. Learns reusable memory guidelines
from trajectories with conflict resolution, feeds them into later sessions,
and reports AppWorld reliability gains with provenance-aware storage.
Targets: Memory.
- EvolveBank — Distills success and
failure trajectories into a deduplicated strategy bank, tracks downstream
win rates, freezes the bank for held-out τ-bench evaluation, and reports a
parity result.
Targets: Memory.
- ExpeL —
Paper. Distills transferable insights
from successes and failures without updating model weights.
Targets: Memory.
- GeoForge — Converts grounded
trajectories into workflow-graph, action-experience, and SOP memories, then
safety-gated distillation reuses them for tool planning across geospatial
benchmarks.
Targets: Memory, Skills, Topology.
- KOPE: Experience-Driven Workflow and Experience Graph Memory — Records hardware-kernel
optimization decisions and correctness/performance feedback in an experience
graph, then retrieves it under a fixed budget for continual optimization.
Targets: Memory.
- Living-Harness — Converts evaluated
trajectories into bounded episodic-memory and state-graph repairs that
persist across episodes and improve interactive benchmark performance.
Targets: Memory, Topology.
- Luclas — Learns from task outcomes
through persistent SQLite episodes and lessons plus versioned self-updating
policies, with daily compression, explicit corrections, and inspectable
drift safeguards.
Targets: Memory, Topology.
- Membrane —
Paper. Evolves a contrastive
safety-memory store from paired harmful/safe prompts and label-free
test-time review, with HarmBench and AgentHarm evaluations.
Targets: Memory.
- MemOS — Applies natural-language
feedback and correction to persistent memory, adds tiered skill evolution,
and publishes LoCoMo, LongMemEval, and OmniMemEval results.
Targets: Memory, Skills, Knowledge.
- MemSkill —
Paper. Learns memory-skill selection and
evolves reusable routines from difficult cases.
Targets: Memory, Skills.
- MetaMem —
Paper.
Edits retained memory-use experiences from judged answers and reflections,
then evaluates selected checkpoints on held-out folds of long-term memory QA.
Targets: Memory.
- OpenViking — Stores memories,
resources, and skills in a browsable context filesystem; commits session
experience to long-term memory and reports LoCoMo/tau2-bench gains with
reproducible scripts.
Targets: Memory, Knowledge, Skills.
- R2-MAD: Remember and Reweight —
Paper. Builds persistent debate
experience memory, retrieves cases using consensus-aware states, and
reweights peer influence by historical reliability across four reasoning
benchmarks.
Targets: Memory.
- Recuris —
Paper. Evolves targeted Skill Memory
from structured traces with paired held-out validation, releases frozen
splits, and reports cross-task transfer across 37 model–benchmark pairs.
Targets: Memory, Skills.
- Reflexio — Turns user corrections
and successful interactions into persistent profiles and playbooks,
aggregates approved cross-user lessons, and reports a warm-baseline GDPVal
comparison across two host agents.
Targets: Memory, Skills.
- Reflexion —
Paper. Stores verbal reflections in
episodic memory for later trials.
Targets: Memory.
- ReMe —
Code.
Distills execution experience into reusable memory, adds successful lessons,
and prunes low-utility records; compares fixed and dynamically updated pools
on tool-use benchmarks.
Targets: Memory.
- RoMeRL —
Paper. Uses bounded task-slot memory
states whose contents and utilities update from environment rewards,
reducing persistent reward contamination across lifelong benchmarks.
Targets: Memory.
- Rudder — Preserves reviewed
lessons, decisions, and skills from agent work in durable context for future
runs, with a documented local-case-equal GDPVal harness comparison.
Targets: Memory, Skills, Topology.
- SelfMem — Uses memory tools and feedback
signals to evaluate and refine a reusable long-context memory strategy.
Targets: Memory.
- SimSkill —
Paper. Explores SUMO traffic tasks,
verifies solutions with action–critic feedback, and consolidates episodic,
procedural, and semantic memories into a reusable library evaluated on two
held-out benchmarks.
Targets: Memory, Knowledge, Skills.
- SSE-Bio —
Paper. Maintains structured reasoning
and template memory, trains a retrieval proxy from decision-contrastive
feedback, and reports gains on three biomedical QA benchmarks.
Targets: Memory, Parameters.
- Tree-of-Experience — Organizes reusable
experience as a reasoning-aligned tree, calibrates path reliability from
environmental outcomes, and evaluates transfer on Game of 24 and
FinEvolveBench.
Targets: Memory.
Knowledge
- Adaptive Reflective Interactive Agent (ARIA) — Maintains a
timestamped knowledge repository from targeted human guidance, detects
conflicts, and evaluates adaptation on changing-domain tasks.
Targets: Knowledge.
- CoEvoKG —
Paper. Writes verified search evidence
back into a knowledge graph that co-evolves with a proposer–solver loop and
is evaluated on six multi-hop QA benchmarks.
Targets: Knowledge, Co-evolution.
- ContDa —
Paper. Rewrites a persistent
tool-documentation corpus from API observations and relation-aware exploration,
evaluating adaptation and retention on evolving StableToolBench and RestBench toolsets.
Targets: Knowledge.
- Knowledge-Centric Self-Improvement —
Paper. Runs disposable agents that
distill evidence-grounded forums into shared knowledge, then seeds later
attempts; gains transfer across held-out tasks and model families.
Targets: Knowledge.
- ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving — Evolves Lean-verified proof
DAGs with neural variation operators, retaining checked schemas across
problems and exposing residual subgoals for typed recombination across three
benchmarks.
Targets: Knowledge, Topology.
- VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following — Synthesizes multimodal
instructions from persistent constraint memory, writes verifier and
target-model failures back into later rounds, and improves MM-IFEval while
preserving general capability across seven public benchmarks.
Targets: Knowledge, Memory.
Skills
- AgentDescent —
Paper.
Evolves skills, prompts, harness modules, and verifiers through parallel
proposals, reward scoring, and asynchronous aggregator merges, with
live-model results and per-run raw data.
Targets: Skills, Tools, Topology.
- cambium — Admission-gates evolving
skill, tool, and prompt libraries, measures retrieval separately, rejects
reward-hacking candidates, and reports held-out transfer from a reproducible
27-task demo.
Targets: Skills, Tools, Topology.
- Coalition-Aware Skill Reliability —
Audits coalition-level skill contributions and masks transfer-harmful
entries, improving LoCoMo, LongMemEval, HotpotQA, and ALFWorld while
exposing isolation-evaluation failures.
Targets: Skills.
- COBRA-Skills — Allocates skill
evaluations with a learned reward predictor and replaces low-scoring skills
through mutation and crossover, with separate optimization and held-out test
stages.
Targets: Skills, Co-evolution.
- CoEvoSkills —
Paper. Builds multi-file Skills through
a generate–verify–refine loop, co-evolving a skill generator and surrogate
verifier before isolated fresh-agent transfer tests.
Targets: Skills, Co-evolution.
- CoSkill —
Paper. Jointly trains reasoning and
meta-skill agents over a hierarchical Skill Bank, promoting lineage-tracked
edits from improvement rewards with ALFWorld/WebShop results.
Targets: Skills, Parameters.
- ERSkill — Evolves reusable retrieval
programs and a learned router from judged rollouts, retaining validation-selected
skills for memory QA and cross-dataset transfer. Official code not located;
reproduction unverified.
Targets: Skills, Parameters, Memory.
- Evo-Harness
— Paper. Compiles noisy one-shot
trajectories into reusable skill harnesses for cross-task adaptation,
evaluating a frozen agent across five realistic benchmarks.
Targets: Skills, Topology.
- From Memory to Skills — Governs the
evidence-grounded conversion of retained traces into callable skills.
Targets: Skills, Memory.
- HypoForge — Learns reusable scientific
skills from stage-specific critique and execution outcomes, enabling
continual improvement without fine-tuning and outperforming framework and
skill-level baselines.
Targets: Skills.
- Learning Globally Reusable Skills for Coding Agents — Co-evolves a skill-relation
graph with global skill consolidation and replay verification, reporting
cross-task and regression-aware gains on coding-agent tasks.
Targets: Skills, Topology.
- Meta Context Engineering (MCE) —
Paper. Co-evolves context-engineering
skills and context artifacts through agentic crossover and bi-level
optimization, with released artifacts and five-domain offline/online
evaluation.
Targets: Skills, Knowledge.
- MUSE-Autoskill — Treats skills as
testable, reusable assets that are refined across tasks.
Targets: Skills.
- OpenSkill —
Paper. Builds and refines reusable
skills against self-created, evidence-grounded virtual tests, then evaluates
frozen skills on held-out target tasks.
Targets: Skills.
- PenguinHarness — Agents
evaluate and optimize their own Skills from benchmark feedback, snapshot
each round, and expose observable traces for review.
Targets: Skills.
- PILOT in the Loop — Steers active
workers during execution and distills procedures and failure modes into
reusable skills and memory, improving three long-horizon benchmarks.
Targets: Skills, Memory.
- PRACTICE — Trains a dedicated learner to
add, refine, merge, and remove a persistent skill library from trajectories,
improving successive rounds for frozen embodied executors.
Targets: Skills.
- RedEvoAgent — Distills cross-case
red-team trajectories into attack skills, retains only validation-improving
updates through a ratchet, and transfers across target harnesses. Run only
in an isolated sandbox.
Targets: Skills.
- RethinkSkill —
Paper. Provides a reproducible
controlled skill-evolution harness that compares success-only, failure-only,
and mixed feedback across 42 matched runs with held-out, robustness, and
transfer checks.
Targets: Skills.
- Retrospective Harness Optimization —
Paper. Uses unlabeled past trajectories
and self-preference to retain skill and tool updates that improve held-out
behavior.
Targets: Skills, Tools.
- RewardHarness —
Paper. Evolves scoring skills and tool
prompts from preference feedback, retaining updates through held-out
validation and rollback.
Targets: Skills, Tools.
- SAGE — Accumulates a
persistent skill library and trains skill generation and use with
outcome-grounded rewards.
Targets: Skills, Parameters.
- Search2Skill — Searches for capability
gaps, distills external evidence into a persistent skill library with
rubric-based reinforcement learning, and improves streaming and held-out
expert-domain evaluation.
Targets: Skills, Parameters.
- self-evolve — Runs multi-round
skill or repository self-iteration in an isolated worktree with
deterministic acceptance, heterogeneous signals, archive lineage, and
rollback; its tests cover 555 cases.
Targets: Skills, Tools.
- SESA: Self-Evolving Search Agents —
Paper. Distills informative self-play
failures into bounded skill memory that reshapes the solver and challenger
frontier across held-out QA benchmarks.
Targets: Skills, Memory, Parameters, Co-evolution.
- skill-up — Turns structured
evaluation failures into persistent skill and regression-suite repairs
through repeated skill-upper iterations with rule, script, or agent judges.
Targets: Skills, Co-evolution.
- SkillAdam —
Paper. Refines reusable skill documents using
persistent issue histories and volatility-controlled edits, retains accepted
revisions, and evaluates frozen-agent performance on separate test partitions.
Use isolated environments.
Targets: Skills, Memory.
- SkillGLoW: Procedural-Family Skill Consolidation — Aggregates task-local
skills into procedural-family priors, admits them only after non-degradation
checks, and transfers compact procedures across four continual task domains.
Targets: Skills.
- SkillHEX — Uses falsifiable self-tests
and evidence-guided tree search to explore persistent skill revisions under
sparse feedback, evaluated on 87 SkillsBench tasks.
Targets: Skills.
- SkillHone —
Paper. Evolves whole skill folders
through persistent decision history, practice probes, and regression-gated
PRs, with validation-gated deep-research evaluation. Run only in an isolated
sandbox.
Targets: Skills.
- SkillOpt —
Paper. Optimizes natural-language
procedures from scored trajectories with validation-gated updates.
Targets: Skills.
- SkillProx —
Paper. Runs a closed-loop
forward/backward skill-evolution process: measured outcomes roll back
regressions, and utility audits gate consolidation or removal before
held-out evaluation; official code is forthcoming.
Targets: Skills.
- SkillZip — Compresses evolving skills
with typed structural sharing and coverage constraints; its Zip-on-Write
mode incorporates each patch without replaying tasks or reparsing history.
Targets: Skills.
- SkillZip Pro — Continually compresses
progressively loaded skill bundles after each evolution patch, preserving
routes and reporting 38% bundle-token reduction without quality loss in a
production moderation harness.
Targets: Skills.
- TRACE —
Paper. Refines a persistent Skill Bank
by contrasting successful and failed trajectories, then reports Pass^3 gains
on public and hidden CAR-bench sets.
Targets: Skills.
- Voyager —
Paper. Grows an executable skill library
through environment feedback and an automatic curriculum.
Targets: Skills, Co-evolution.
- When Self-Evolution Backfires — Shows
skill pools can contaminate later skills and proposes pre-commit
heterogeneous critics plus marginal-gain subset selection, evaluated on
Terminal-Bench 2.
Targets: Skills.
- WikiSkill — Consolidates execution
experience into a persistent wiki that guides later skill updates, with
cross-model transfer and ablations showing knowledge accumulation matters.
Targets: Skills, Knowledge.
- xskill —
Paper.
Distills anonymized agent trajectories into versioned team skills,
canary-tests revisions on real traffic, and reports benchmark gains with
rollback-ready lineage.
Targets: Skills, Memory.
- AgentFactory —
Paper. Stores successful
solutions as executable Python subagents, refines them from execution
feedback, and evaluates reusable capability growth across later tasks. Run
only in an isolated sandbox.
Targets: Tools.
- Mem²Evolve —
Paper. Couples experience distillation
with dynamic creation of tools and expert agents.
Targets: Tools, Memory, Skills, Co-evolution.
- SciToolAgent-Evo — Evolves an
ontology-backed scientific tool graph and skill/experience memory through
contrastive trajectories and bandit-gated acquisition, evaluated on 900
OpenSciToolBench tasks.
Targets: Tools, Knowledge, Memory.
Topology
- A-Evolve —
Paper. A position paper and framework
treating deployment-time improvement as optimization over persistent agent
state; empirical scope requires separate extraction.
Targets: Topology.
- A-SR: Self-Evolving Agentic LLMs for Symbolic Regression — Routes evaluator feedback
through role policies and process memory, adapts coordination within runs,
and distills trajectories across runs with held-out symbolic-regression
evaluation.
Targets: Topology, Memory, Parameters.
- ADAS —
Paper. Searches agent designs expressed
as code and retains an archive of evaluated candidates.
Targets: Topology.
- AegisEvo — Governs sandboxed harness
search with statistical quality, safety, canary, promotion, and rollback
gates, publishing deterministic fixtures and reproducible reports.
Targets: Topology.
- AgentSquare —
Paper. Searches a modular space of
planning, reasoning, tool-use, and memory components.
Targets: Topology, Memory, Skills, Tools.
- AutoDesign —
Paper. Learns a reusable design harness
around fixed models, retaining one-component updates only when training
improves without regressing an independent development set, then evaluates
matched configurations on PosterBench.
Targets: Topology, Tools, Skills.
- CausalForge —
Paper. Uses Lean-checked proofs and
statement audits in a self-improving theorem pipeline, with public formal
libraries and run records.
Targets: Topology.
- CineForge: Self-Improving Agents for Long-Horizon Video Generation — Consolidates production
trajectories into bounded, stage-local policy patches, validates them by
structural replay and paired evaluation, and improves scores on a 100-script
suite plus two public benchmarks.
Targets: Topology.
- Darwin Gödel Machine —
Paper ·
Article. Rewrites coding-agent implementations and
empirically validates descendants. Run only in an isolated sandbox.
Targets: Topology, Skills, Tools.
- EMAS: Evolving Multi-Agent System —
Paper. Converts recurring trace
diagnoses into prompt or topology revisions, paired-validates each
candidate, and persists accepted versions while retaining rejected proposals
for audit.
Targets: Topology.
- EvoAgentX — Builds, evaluates, and
evolves multi-agent workflows with pluggable optimization algorithms.
Targets: Topology.
- Factory — Provides
tracker-controlled coding-agent loops with isolated worktrees, repeatable
verification, CI/review gates, Git lineage, and offline regression demos.
Run only in an isolated sandbox.
Targets: Topology.
- From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis — Converts successful and failed
RCA trajectories into atomic expertise updates, admits them through
dual-gate verification, and reports gains on two public benchmarks plus an
industrial deployment.
Targets: Topology, Knowledge.
- GPTSwarm —
Paper. Represents language agents as
graphs and optimizes node prompts and connectivity.
Targets: Topology.
- HarnessEvolve — Aligns failed runs with
reference trajectories, then quality- and performance-gates harness
snapshots against held-out validation to reduce shortcut learning and
forgetting.
Targets: Topology.
- HarnessLens —
Paper. Evolves OpenCode, Codex CLI, and
Pi harnesses through behavior-aware diagnosis and selective verification,
with blind-test entrypoints, pinned reproducibility, and four-benchmark
evaluation. Run only in an isolated sandbox.
Targets: Topology.
- JIT-Agent —
Paper. Distills performance signals from
an archive of prior harness configurations to synthesize, repair, and
improve task-adaptive harnesses across model families and benchmarks.
Targets: Topology.
- MANTA: Multi-Agent Network Topology Adaptation — Adapts roles, links,
execution order, and visibility from trace audits during execution, retains
a cross-run topology playbook, and evaluates transfer across five
benchmarks.
Targets: Topology, Memory.
- Meta-Harness — Implements
recorded-tape harness replay, persistent checkpoints, candidate selection,
and separate search/holdout protocols; replay does not imply fresh-model
reproducibility. Run only in an isolated sandbox.
Targets: Topology, Skills.
- Meta^n: Recursive Self-Improvement through Emergent Depth —
Paper. Recursively writes executable
helpers and pre-processors from lower-layer traces, archives evaluated
chains, and reports gains across eight benchmark families. Run only in an
isolated sandbox.
Targets: Topology, Skills.
- MetaVideoAgent —
Paper. Profiles a video distribution,
diagnoses trajectory failures, and evolves responsible modules across four
rounds with evolution and held-out VA-EvoBench splits. Run only in an
isolated sandbox.
Targets: Topology.
- Naive Prompt Optimization — Iteratively
revises prompts from teacher-model rollout feedback and transfers
single-lineage improvements across tasks and interactive games.
Targets: Topology.
- Open-Ended Optimization (OEO) — Composes
the improvement route online under fixed objectives, budgets, data
boundaries, and evaluators, comparing persistent skill updates with SkillOpt
and GEPA.
Targets: Topology, Skills.
- OpenEvolve —
Provides an open evolutionary coding loop inspired by AlphaEvolve.
Targets: Topology.
- Proteus — Provides a
harness-agnostic, snapshot-based evolution loop with evaluator gates,
crystallization tests, rollback, and git histories for measuring persistent
change. Run only in an isolated sandbox.
Targets: Topology.
- Raven — Combines terminal execution,
tracing, durable memory, skills, reusable workflows, and an opt-in Evolver
with independent benchmark harnesses for long-running agent improvement. Run
only in an isolated sandbox.
Targets: Topology, Memory, Skills.
- Reef — Connects live
inference, feedback, candidate training, evaluator selection, and versioned
delivery for model weights or harness skills, with recipes and task-level
result reports. Run only in an isolated sandbox.
Targets: Topology, Parameters, Skills.
- ROSClaw — Provides a simulation-first
physical-agent control plane that turns feedback into hashed controller
candidates, fail-closed gates, and rollback targets; real-hardware
activation is blocked. Run only in an isolated sandbox.
Targets: Topology.
- RSIHub — Runs evaluator-driven
evolution over prompts, skills, harnesses, and agent code with frozen
scoring, bounded mutation, Git lineage, rollback, and reproducible benchmark
recipes. Run only in an isolated sandbox.
Targets: Topology, Skills.
- RubSE —
Paper. Generates UI code with typed
rubrics as visual feedback, retains selected repair history across rounds,
and evaluates six VLMs on three UI-to-code benchmarks for stable iterative
improvement.
Targets: Topology, Memory.
- Self-Improving Agent Ecosystem
— Provides contracts, schemas, deterministic fixtures, and validation for
evaluator-driven loops with isolated candidates, evidence lineage, promotion
gates, rollback, and external health checks.
Targets: Topology.
- VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding — Searches executable
context constructors around a frozen VLM, retaining evaluation-selected
variants and transferring the selected harness to additional long-video
benchmarks.
Targets: Topology.
- yoyo — Runs an autonomous
coding-agent loop that reads source and community issues, test-gates commits
or reverts, and synthesizes durable memory across sessions. Run only in an
isolated sandbox.
Targets: Topology, Memory.
Co-evolution
- Agent0 —
Paper. Co-evolves a task curriculum and
a tool-using executor without human-curated task data.
Targets: Co-evolution, Parameters.
- CORAL —
Paper. Runs autonomous coding-agent
organizations in isolated worktrees, sharing persistent attempts, notes, and
skills while a grader scores commits and agents iterate on open-ended tasks.
Targets: Co-evolution, Topology, Tools.
- Double Ratchet
— Paper. Co-evolves an inspectable
drawback-detector metric with a lifecycle-managed skill library, using
anchored references and locked-set validation to expose evaluator collapse.
Targets: Co-evolution, Skills.
- Environment Evolution for Terminal Agents — Evolves terminal environments
off-policy along difficulty directions to sustain learning signals,
improving long-horizon RL on Terminal-Bench 2.1.
Targets: Co-evolution.
- Group-Evolving Agents (GEA) —
Paper. Evolves populations of agent
workflows and tools through archived experience sharing, retaining
cross-lineage improvements on held-out coding benchmarks. Run only in an
isolated sandbox.
Targets: Co-evolution, Topology, Tools.
- HELIX —
Paper. Composes source-traceable harness
variants, evaluates sibling rollouts, and turns verified successes,
regressions, and preferences into data for later model updates; reports
LiveCodeBench and SWE-Bench results.
Targets: Co-evolution, Topology.
- J-Zero — Co-evolves challenger, solver,
and judge from zero data, using known-order preference pairs and adversarial
tasks to improve through ten iterations.
Targets: Co-evolution, Parameters.
- MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph — Distills sessions into durable
wisdom assets and a typed graph, then feeds controlled operational evidence
back into workflow optimization and the evolving curation strategy.
Targets: Co-evolution, Knowledge, Topology.
- SafeEvolve —
Paper. Turns on-policy safety
trajectories into bounded safety-prompt and SkillBank revisions, then trains
policy use with SFT and GRPO under safety and utility gates. Run only in an
isolated sandbox.
Targets: Co-evolution, Parameters, Skills.
- SBCO: Self-Supervised, Verifier-Grounded Harness Optimization — Jointly updates a
decomposed verifier bank and planning-agent harness from self-graded
feedback via block-coordinate ascent, without human labels.
Targets: Co-evolution, Topology.
- Self-Modifying Lean Proof Agents —
Co-evolves a Lean proof workflow and benchmark curriculum, using compiler
and Lean verification feedback across 15 generations with a held-out miniF2F
split.
Targets: Co-evolution, Topology, Tools.
- Task-CoEvolve — Co-evolves harness code
and validation-task sampling, focusing evaluation on discriminative tasks
while retaining unbiased full-set estimates and reducing evaluation cost.
Targets: Co-evolution, Topology.
- WebEvolver
— Paper. Co-trains a policy
and world model from real trajectories, then uses the model for synthetic
training and look-ahead evaluation.
Targets: Co-evolution, Parameters.
Co-evolution denotes coupled change, not simply multi-agent execution. See
boundary decisions before assigning secondary targets.
Benchmarks and Evaluation
Evaluate retained improvement separately from extra inference compute. The
subsections distinguish task streams, measurement controls, safety audits, and
general task benchmarks; none automatically demonstrates persistent learning.
Longitudinal and adaptation evaluation
- AgentStream —
Paper. Evaluates five self-evolving
methods in isolated, sequential, and interleaved task streams across three
models, exposing model-, method-, and stream-dependent reliability.
- AI4AI-Bench — Releases 10 frozen
research repositories, hidden evaluators, 29 configurations, and every
scored submission for testing whether agents can rewrite training
algorithms.
- ASPIRE — Hides expert-authored tasks
behind vague goals and scores whether agents choose data, updates, and
validation signals, exposing transfer and stability failures.
- Continual Skill Bench —
Paper. Evaluates continual skill reuse
across five domains and 100-task sequences, comparing in-context adaptation
with explicit skill maintenance.
- Evo-Bench —
Paper. Holds policy, seed harness, and
budget fixed while scoring autonomous harness evolution on disjoint
validation and evaluation suites.
- EvoHarnessBench — Separates fresh-state
deployment from persistent adaptation under externally expanding tool, skill,
and specialist-agent pools, measuring retention and transfer with held-out evaluation.
- Experience-driven Lifelong Learning —
Proposes a framework and benchmark for continuous agent growth.
- FinEvo-Bench — A longitudinal benchmark
with 120 financial workflow tasks, interleaved streams, paired state-reset
controls, expert-validated rubrics, and compliance metrics for retained-
experience gains.
- PACE-Bench —
Paper. Tests whether code-driven designs
adapt after source-to-target physics mutations, with diagnostic sandbox
feedback across 144 pairs and 180 evaluation environments.
- reef-eval — Provides
Harbor-compatible autoresearch and continual task streams with state
snapshots, tamper-resistant scoring, and anytime, forgetting, and transfer
metrics.
- RSIBench-Data —
Paper. Audits whether researcher agents
turn checkpoint feedback into reusable training-data strategies under fixed
training, evaluation, and budget controls.
- S3Gym — Separates permissive exploration
from held-out game evaluation while comparing history, summary-memory, and
parameter updates for self-improvement.
- SEAGym: An Evaluation Environment for Self-Evolving LLM Agents —
Paper. Measures persistent harness
updates across train, frozen validation, held-out ID/OOD transfer, replay,
and cost views, with checkpointed states and reproducible configurations.
- StudyBench —
Paper. Releases fixed textbook
materials, application/transfer splits, and evaluation scripts to measure
whether self-evolution converts study material into transferable capability.
- tide-eval — An
Apache-2.0 evaluation infrastructure for self-evolving agents on Harbor:
carries memory, skills, or harness state across task streams with per-step
snapshots, and reports anytime, forgetting, and transfer metrics.
Reliability and evaluation controls
- Cheap Verifiers, Large Blind Spots —
Audits self-improving verifier cascades with true-error controls, showing
in-loop metrics can improve while delivered quality degrades under
blind-spot feedback.
- On the Fragility of Self-Improving Agents —
Paper. Re-evaluates memory-based
self-improving agents across repeated runs and shuffled task orders,
releasing trajectories to measure variance, order sensitivity, and
specification effects.
- Phantom Gains — Audits self-improvement
claims against measured frozen controls, replacing noisy transition
statistics with per-problem tests and false-discovery-rate control.
- Sealed Exogenous Acceptance Loop (SEAL)
— Keeps its audit hidden from evolving policies and self-tests, returning
only accept/reject and retaining the whole incumbent after a clear
regression.
Safety, poisoning, and recoverability
- Auditing Harness Tampering — Audits
authorization, provenance, and completeness violations in self-modifying
harnesses using tampered-benign pairs, localization tasks, and
real-trajectory persistence analysis.
- Auditing Self-Evolution in Financial Agents — Audits three self-evolving
agents with sealed endpoints, state replay, execution-grounded checks, and
security metrics, exposing capability gains that increase exposure and
unauthorized state changes.
- ECLIPSE — Builds and iteratively
verifies stealthy tool-chain injections, releasing LASE-Bench for
long-horizon safety evaluation. Run only in an isolated sandbox.
- EVOMAL: Self-Poisoning in Self-Evolving Coding Agents — Measures self-poisoning
propagation in evolving skill libraries across six models and 153 SWE-bench
tasks, and evaluates a counter-prompt defense without task-completion loss.
- EvoSkill Injection — Defines a red-team
threat model, EvoSkillBench trajectories, and post-attack safety tests for
persistent malicious-skill formation and repeated activation. Run only in an
isolated sandbox.
- EvoUndo — Evaluates recoverability of
model-generated prompt, tool, middleware, and harness mutations across
counterfactual states on 600 unseen tasks, exposing grounding and
recovery-language bottlenecks.
- HarnessRisk —
Paper. Benchmarks agent-harness safety
across configuration, capability extension, runtime, persistence, action
control, and recovery using sandboxed cases with trajectory evidence and
explicit utility, attack, persistence, and detection metrics.
- PerMemSafe —
Paper. Benchmarks
implicit personalized safety across evolving long-horizon memory, with
committed data, evaluation logs, and a risk-aware memory baseline.
- SIR: Self-Improving Red Teaming —
Distills failed computer-use attacks into reusable principles with
deterministic filesystem, service, and permission oracles; report transfer
evidence and run only in an isolated sandbox.
- SkillJack
— Paper. Releases poisoned-trajectory
data and cross-system experiments showing experience-to-skill backdoors
survive source deletion and can activate on benign queries. Run only in an
isolated sandbox.
- When Experience Becomes Instruction —
Demonstrates trajectory-poisoning attacks that promote malicious behaviors
into persistent skill banks, with transfer tests across six LLM evolvers and
two architectures.
General agent benchmarks
- AgentBench — Evaluates agents across
multiple interactive environments.
- GAIA — Tests assistants on tasks
requiring reasoning, tools, and multimodal information.
- SWE-bench — Supplies real-world
software issues used by self-improving coding-agent systems.
For comparison tables, record the frozen backbone and baseline, optimizer-visible
versus held-out data, update budget, multiple seeds or uncertainty, transfer and
forgetting, safety regressions, and retained artifact versions. Use the
evidence extraction protocol, not a single headline score.
Articles and Technical Posts
Community and Long-Term Outputs
The repository is community-first and is building a structured evidence base.
The evidence may support a living survey and conference tutorial after the
public maturity gates in the Roadmap are met.
Repository contributions are credited publicly. Sustained, substantive work in
curation, reproducibility, analysis, writing, software, visualization, or
revision can qualify contributors for survey-paper authorship. Decisions follow
transparent CRediT-style roles rather than maintainer status.
Contributing
Suggestions and pull requests are welcome. Read
CONTRIBUTING.md before submitting a resource.
This index favors evidence over hype. A large star count is neither necessary
nor sufficient; a resource should make the update target, feedback loop,
persistent artifact, and evaluation legible.
License
MIT