tfatykhov/awesome-agent-memory

Curated research on memory systems for LLM agents

23

25 commits

updated Sep 27, 2026

See the code

README

🧠 Awesome Agent Memory

Awesome

A curated collection of research papers on memory systems for LLM-based agents — covering architecture, retrieval, forgetting, consolidation, evaluation, the cognitive science that inspires it all, and the neuromorphic hardware that may one day implement it.

Why this list? Every agent framework bolts on a vector store and calls it "memory." These papers show what memory actually requires: admission control, consolidation loops, forgetting mechanisms, typed multi-store architectures, and retrieval strategies that go far beyond cosine similarity.

Maintained by @tfatykhov · Built alongside Nous, a cognitive AI agent with typed multi-store memory.

🚀 Start Here

If you want to...Start with
Understand the landscapeMemory Survey (taxonomy) → Anatomy (critical analysis)
Build agent memoryxMemory (why RAG isn't enough) → A-MAC (admission control) → Mem0 (production system)
Add forgetting/consolidationSleepGate → CraniMem
Evaluate your memory systemStructMemEval → Anatomy (why benchmarks are broken)
Train memory with RLAgeMem (tool-based ops) → EMPO² (exploration) → MemFactory (framework)
Understand memory privacy riskADAM (extraction attack) → FSFM (forgetting as defense)
Unify memory ↔ skills ↔ rulesExperience Compression Spectrum → Externalization
Make memory writes safe & reversibleMemTX (transactional belief commit) → ChronoMem (versioning + rollback)
Ship something this weekImplementations & Reference Code (curated repos, mechanism-by-mechanism)
Explore the neuroscienceNeuromorphic section

Contents

Verification Legend

SymbolMeaning
✅Venue confirmed — verified via official proceedings, OpenReview, or arXiv metadata
📄Self-reported — metrics from authors' own evaluation; exercise caution
⚠️Unverified — plausible claim but not independently confirmed
🔬Editor's synthesis — cross-paper connection identified by the maintainer, not claimed by original authors

Surveys & Taxonomies

PaperAuthorsDateKey Contribution
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory ProblemTangAug 2026The first serious attempt to say what a memory is, formally: memory is a basis, knowledge is its span, and answerability is a coverage problem — a query is answerable exactly when some single item in the span covers it. Optimal memory becomes the capacity-constrained maximizer of expected coverage, tracing a utility–capacity frontier that serves as a common yardstick for comparing memory systems. Also treats noise (coverage vs precision when the store holds false claims) and formalizes continual memory as a sequential MDP — memory is the state, writing is the action, query-time utility is the delayed reward. Instantiated concretely on Homer's Odyssey. ⚠️ Preprint; solo-author.
Memory in the Age of AI Agents: A Survey47 authorsDec 2025Unified taxonomy across 3 lenses: Forms (token/parametric/latent), Functions (factual/experiential/working), Dynamics (formation/evolution/retrieval). Distinguishes agent memory from RAG and context engineering. Hugging Face Daily Paper #1 ⚠️ (unverifiable — HF archive page returns 404).
Anatomy of Agentic MemoryJiang, Li, Wei et al.Feb 2026Structured taxonomy of Memory-Augmented Generation (MAG) systems. Critically analyzes empirical fragility: benchmark saturation, metric misalignment, backbone-dependent variance, overlooked latency costs. Evaluates 5 systems (LOCOMO, A-Mem, MemoryOS, Nemori, MAGMA).
The Landscape of Agentic RL for LLMsGuibin Zhang et al. (Oxford, Shanghai AI Lab, NUS, UIUC, UCL)Sep 2025Synthesizes 500+ works. Reframes LLMs as autonomous agents using POMDPs. Covers planning, tool use, memory, reasoning, self-improvement.
From Static Templates to Dynamic Runtime GraphsYue et al. (IBM Research)Mar 2026Agentic Computation Graphs (ACGs) framework — distinguishes workflow templates, realized graphs, and execution traces. Organizes ~40 papers by when structure is determined.
Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging FrontiersPengfei DuMar 2026Write–manage–read loop formalization. 3D taxonomy: temporal scope × representational substrate × control policy. Five mechanism families: context-resident compression, retrieval-augmented stores, reflective self-improvement, hierarchical virtual context, policy-learned management. Proposes vision of a "foundation model for memory control."
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness EngineeringZhou, Chai, Chen, Guo, Shan, Song, Xu et al.Apr 2026Unified review across four externalized components: memory stores, reusable skills, interaction protocols, runtime harness. Argues modern agent capability comes from reorganizing the runtime around the model, not from changing weights. Connects memory literature with skill-discovery and protocol-design literatures that rarely cite each other.
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM AgentsZhang, Wang, Cui, Qiu, Li, Zhu, HeApr 2026Positions memory/skills/rules on a single compression axis (5–20× episodic, 50–500× procedural, 1000×+ declarative). Citation analysis of 1,136 references finds <1% cross-community citation between memory and skills literatures. Identifies the "missing diagonal" — no current system supports adaptive cross-level compression.
From Storage to Experience: A Survey on the Evolution of LLM Agent Memory MechanismsLuo, Tian, Cao, Luo et al.May 2026ACL 2026 Findings ✅ — Maps the field's evolution from passive storage stores toward experience-centric, self-evolving memory. Companion paper-list: Evolving-LLM-Agent-Memory-Survey.
Graph-based Agent Memory: Taxonomy, Techniques, and ApplicationsYang, Zhou, Xiao, Dong et al.Feb 2026First focused taxonomy of graph-based agent memory: node/edge designs, construction strategies (entity-centric vs event-centric vs hybrid), retrieval (one-hop vs multi-hop vs subgraph), and update/forget operators. Useful companion to A-MEM, HeLa-Mem, GAM.
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent MemoryOrogat, MansourMay 2026Argues record-level correctness (rows, embeddings, edges) cannot satisfy long-term memory's needs, causing four recurring failure modes: unregulated growth, missing semantic revision, capacity-driven forgetting, read-only retrieval. Formalizes Governed Evolving Memory (GEM) — state-level operators (ingestion, revision, forgetting, retrieval) with six correctness conditions on the trajectory, not individual records. Prototype MemState on a property-graph backend. Proves no record-level system can satisfy the conditions regardless of storage model. ⚠️ Preprint.

Memory Architectures

PaperAuthorsDateKey Contribution
Metis: Memory Foundation ModelZhang, Guo, Sun, Zhang, Hao, Lin et al. (17 authors)Jul 2026Challenges this list's prevailing external-module framing. Argues agent memory should be native to the backbone, and formalizes native memory as (a) a persistent, dynamically evolving memory state inside the model and (b) native memory procedures that store and use information through model computation. Metis equips a foundation model with a native memory state accessed via memory attention, acquired through mid-training on large-scale memory-specific data. Online memory maintenance is gradient-free — a memory update requires only a forward pass, and all learned weights remain frozen at inference. Project and model checkpoints released. ⚠️ Preprint.
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and SustainabilityZhou, Yu, Wei, Wu, Ouyang, Jiao et al. (11 authors)Jul 2026The first systematic study of the memory design that is actually deployed by default — a directory tree of markdown files the agent itself reads, writes, and reorganizes with generic file tools — which prior research largely passed over. Formalizes it as three roles around one memory filesystem (management, search, execution agents), unifying declarative memory and skills in a single store. The default's two working assumptions do not hold up: what organization reliably buys is search economy (organized stores roughly halve retrieval cost where material is large), not better answers — no agent measured converts organization into accuracy, and organization erodes as the store grows for all but the strongest management agent. Changing the tool set alone reshapes the store as strongly as swapping the model 📄. ⚠️ Preprint.
MAGE: Memory as Agent-Guided ExplorationChen, Lai, Feng, Han, Zhang, Lu, Li et al.Jun 2026Argues semantic-similarity organization mismatches execution-state dependencies on long-horizon tasks — it fragments decision trajectories and mixes valid/erroneous traces. Stores interactions in a hierarchical state tree; agent state derives from the active root-to-current path (subgoal summaries + recent traces + branch hints). Four coupled ops: Grow, Compress, Maintain, Revise. +7.8–20.4pp task success on MemoryArena, -55.1% tokens 📄.
Mechanistic Attention Guidance for Agent Memory Refinement (AGMR)Hong, Qiu, Wang, YaoJul 2026Existing self-evolving memory only inspects textual outputs (trajectories, reflections), risking unreliable error attribution and hallucinated edits. Uses retrieval-head attention as a mechanistic signal — aggregates attention over memory segments × decision steps into a context-utilization matrix that reveals actual memory-use patterns. AGMR corrects/enhances memory on failure, simplifies it on success, and re-executes to verify each update. Beats text-only refinement baselines on interactive decision-making benchmarks. Code released. ⚠️ Preprint.
Supra Cognitive Modes (SCM): A Routed Architecture for Agent MemoryTobkin, YangJul 2026Routes each query — factual lookup, relation-chain/current-state reasoning, or broad long-history synthesis — to a matched retrieval+synthesis payload over one shared ingest substrate (multi-granularity embeddings, extracted triples, fact-version metadata). A frozen semantic classifier + runtime gates dispatch among fused lexical/dense lookup, graph/multi-hop, and stratified long-form synthesis. Reports 84.87% LoCoMo factoid / 68.61% adversarial abstention, 61.49% MAB, 86.00% LongMemEval — but the authors themselves flag causal routing effects and efficiency gains as outside the available evidence 📄 (unusually candid self-limitation disclosure). ⚠️ Preprint.
A-MEM: Agentic Memory for LLM AgentsXu, Liang, Mei, Gao, Tan, ZhangFeb 2025Dynamic memory organization inspired by Zettelkasten. Creates interconnected "memory notes" with content, metadata, and explicit links. Agent autonomously decides when to create, update, or link memories. NeurIPS 2025 ✅.
SYNAPSE: Episodic-Semantic Memory via Spreading ActivationJiang, Chen, Pan et al.Jan 2026Memory as a dynamic graph where relevance emerges from spreading activation rather than pre-computed vector links. Lateral inhibition and temporal decay highlight relevant sub-graphs. Triple Hybrid retrieval.
MAGMA: Multi-Graph Agentic Memory ArchitectureJiang, Li, Li, LiJan 2026Multi-graph based architecture separating different memory concerns into distinct graph structures for AI agents. ACL 2026 Main ✅.
Mem0: Production-Ready AI Agents with Scalable Long-Term MemoryChhikara, Khant, Aryan, Singh, YadavApr 2025Scalable memory-centric architecture with dynamic extraction, consolidation, and retrieval. Graph-based variant for complex relational structures. 26% improvement over OpenAI memory, 91% lower latency, 90%+ token savings 📄 (self-reported). Evaluated on LOCOMO.
CraniMem: Neurocognitively Motivated Gated & Bounded MemoryMody, Panchal, Kar, Bhowmick, KaraniMar 2026Cranial-inspired dual-store: bounded FIFO episodic buffer + knowledge graph (long-term). Goal-conditioned gating. Utility tagging (Importance + Surprise + Emotion). Beats Mem0 by +57.6% on noisy HotpotQA 📄 (self-reported; "noisy" = authors' own distractor injection, not standard benchmark). ICLR 2026 MemAgents Workshop ✅. Code
AdaMem: Adaptive User-Centric Memory—Mar 2026Working, episodic, persona, and graph memories. Key innovation: question-conditioned retrieval planning — resolves target participant first, builds retrieval route combining semantic + relation-aware graph expansion. SOTA on LoCoMo and PERSONAMEM.
PlugMem: A Task-Agnostic Plugin Memory ModuleYang, Galley, Wang, Gao, Han, Zhai (Microsoft Research, UIUC)Mar 2026Converts raw interactions into propositional (facts) + prescriptive (skills) knowledge, organized in a knowledge-centric memory graph. Knowledge — not entities or chunks — is the unit of memory access. Single plug-and-play module outperforms task-specific designs across 3 diverse benchmarks. Highest information density: more useful info per context token consumed 📄.
Memora: Harmonic Memory RepresentationXia, Zhang, Dixit, Harimurugan, Wang, Ruhle, Sim, Bansal, Rajmohan (Microsoft Research)Feb 2026Proves that RAG and KG memory are special cases of this unified framework. Primary abstractions index concrete memory values; "cue anchors" expand retrieval beyond semantic similarity. New SOTA on both LoCoMo AND LongMemEval 📄. ICML 2026 ✅.
MemMA: Coordinating the Memory Cycle through Multi-Agent ReasoningLin, Zhang, Lu, Liu, Tang, He, Zhang, Wang (Microsoft Research)Mar 2026Multi-agent framework: Meta-Thinker → Memory Manager → Query Reasoner. Backward path innovation: synthesizes probe QA pairs, verifies memory, converts failures into repairs BEFORE finalizing. Plug-and-play — improves 3 different storage backends on LoCoMo. Code
H-Mem: Hybrid Multi-Dimensional MemoryYe, Huang, Chen, Zhang (Rutgers)Mar 2026Organizes memory across time AND topic dimensions simultaneously. Mimics associative + hierarchical properties of human memory. EACL 2026 ✅.
H-MEM: Hierarchical Memory for High-Efficiency Long-Term ReasoningSun et al.Mar 2026Multi-level memory storage with positional index encoding of sub-memory at each layer. Confidence-weighted retrieval — attaches memory weights to provide LLMs with uncertainty reference. Key finding: retrieval is ineffective without structured hierarchical storage. EACL 2026 ✅.
GAM: Hierarchical Graph-based Agentic Memory for LLM AgentsWu, Zhang, Lin, Xu, Xu, Chen, Zou et al.Apr 2026Explicitly decouples encoding from consolidation: an event-progression graph captures stream updates online; integration into the topic-associative network is deferred until a semantic shift is detected. Addresses the tension between stream-based fluidity and structured retention.
HeLa-Mem: Hebbian Learning and Associative Memory for LLM AgentsZhu, Li, Zhang, Liu, YangApr 2026ACL 2026 ✅ — Bio-inspired dual-graph: (1) episodic graph evolves via Hebbian co-activation; (2) semantic store populated via Hebbian Distillation — a Reflective Agent identifies densely-connected hubs and distills them into reusable semantic knowledge. Beats prior SOTA on LoCoMo across 4 categories with fewer context tokens 📄. Code
LightMem: Lightweight LLM Agent Memory with Small Language ModelsZhang, Zhang, Chen, Huang, Zheng et al.Apr 2026ACL 2026 ✅ — SLM-driven memory with strict online/offline separation. STM/MTM/LTM tiers with two-stage retrieval (vector coarse → semantic re-rank). ~2.5 F1 over A-MEM on LoCoMo, 83ms retrieval, 581ms end-to-end 📄. Shows that careful SLM use can replace repeated large-model memory calls.
MemMachine: Ground-Truth-Preserving Memory for Personalized AI AgentsWang, Yu, Love, Zhang, Wong, Scargall, Fan et al.Apr 2026Open-source system integrating short-term, long-term episodic, and profile memory in a ground-truth-preserving pipeline. Targets multi-session degradation in standard RAG.
Omni-SimpleMem: Autoresearch-Guided Lifelong Multimodal MemoryLiu, Ling, Qiu, Liu, Han, Xia, Tu et al.Apr 2026Uses autonomous research-agent search over the design space (architecture × retrieval × prompts × data pipeline) to discover effective lifelong multimodal memory configurations. The first paper to treat memory architecture itself as a search target.
Human-Inspired Memory Architecture for LLM AgentsKerestecioglu, Robsky, Vasters, Sharma, Kesselman (Microsoft)May 2026Biologically-grounded architecture with six cognitive mechanisms: sleep-phase consolidation, interference-based forgetting, engram maturation, reconsolidation on retrieval, entity knowledge graphs, hybrid multi-cue retrieval. Introduces a synthetic calibration methodology that derives all thresholds without benchmark exposure — eliminates a common eval-leakage source. First streaming M-tier LongMemEval eval (475 sessions). Dedup-based consolidation: 97.2% retention precision, 58% store reduction (+21.8 pp) on a 13K-issue VSCode dataset 📄.
MRAgent: Memory is Reconstructed, Not Retrieved — Graph Memory for LLM AgentsJi, Li, Hooi (NUS)Jun 2026ICML 2026 ✅ — Represents memory as a Cue-Tag-Content graph where associative tags bridge fine-grained cues to contents. An active reconstruction mechanism folds LLM reasoning into memory access — iteratively exploring and pruning retrieval paths on accumulated evidence, adapting retrieval to the reasoning context while avoiding combinatorial blowup. Up to 23% over strong baselines on LoCoMo + LongMemEval with substantially lower token/runtime cost 📄. The cleanest articulation of the "reconstruction ≠ retrieval" paradigm.
MemPro: Agentic Memory Systems as Evolvable ProgramsLiu, Wang, Wu, Huang, Tao, Song, Zhou, He (ECNU)May 2026Treats the entire memory construction-retrieval (MCR) pipeline as an evolvable program, not just the memory bank. Maintains a version tree of runnable memory-system implementations; an Evolving Agent selects promising versions, diagnoses recurring failures, and synthesizes improved children via failure-mode-guided edit-debug refinement. Beats static and prompt-level evolving baselines on LongMemEval / LoCoMo / HotpotQA / NarrativeQA within a few iterations, with favorable performance-cost trade-off 📄. Code available. ⚠️ Preprint; venue unconfirmed.
MemDreamer: Hierarchical Graph Memory with Agentic Retrieval for Long Video—Jun 2026Extends agent memory to the multimodal long-video setting: a hierarchical graph memory plus agentic retrieval over hours-long video, exploiting the tree+graph structure for question answering. ⚠️ Preprint; venue unconfirmed.
H-Mem (Yu & Fang): Evolving & Retrieving Agent Memory via a Hybrid StructureYu, Fang, Liu, Ma (CUHK-Shenzhen / Huawei)May 2026Distinct from the EACL H-Mem (Rutgers) and the Sun et al. H-MEM listed above — a third system sharing the name. Hybrid tree + graph: a temporal-semantic tree lets short-term memory evolve into a summarizing long-term store, while a parallel knowledge graph captures entity relationships; retrieval exploits both. SOTA on QA across three agent-memory benchmarks 📄. ⚠️ Preprint; venue unconfirmed.
MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context ManagementLiu, Wu, Liu, Zhao, Liu, Li, Zhang, Wang, Guo, LiuJun 2026Extends agent memory to long-horizon mobile GUI agents. Introduces Context-as-Action (ConAct): context management becomes a first-class action emitted by the same policy that selects UI actions, maintaining three structured fields (folded action history, folded UI state, recent step record) instead of passively appending ReAct-style transcripts. 8B model trained on the 2,956-trajectory MemGUI-3K dataset achieves best open-data 8B performance on MemGUI-Bench and generalizes OOD to MobileWorld. Code/data to be released. ⚠️ Preprint.
Multi-Head Recurrent Memory AgentsLi, Yeh, LiJul 2026Diagnoses why recurrent memory agents degrade at long context: decomposing performance into capture vs retention shows retention collapse is the dominant bottleneck, caused by treating memory as one monolithic text block that every update risks overwriting. Multi-Head Recurrent Memory (MHM) partitions memory into independent heads with a stage-wise select-then-update strategy — only one head updates per step, the rest are structurally shielded. The MHM-LRU instantiation is training-free and lifts RULER-HQA retention at 896K tokens from <30% to 73.96%. Architecture, not model behavior, is the lever. 📄 ⚠️ Preprint.

Memory Admission & Gating

PaperAuthorsDateKey Contribution
When Not to Write Memory: Governing False Promotion from Correlated Agent Traces (GovMem)Qi, Xu, LiJun 2026Reframes admission as a write-path governance problem: repeated observations aren't independent evidence if copied from a shared source, induced by a shared prompt, or valid only in a narrower scope. GovMem estimates dependency-aware support, retrieves counterevidence, assigns scope, and outputs promote / reject / needs-review. Cuts false promotion 0.597→0.040 (synthetic) at 0.960 recall; on a human-labeled real-trace subset, false promotion falls to 0.032 but held-out false promotion stays 0.111 at 0.692 review burden. On 133 high-impact external coding-agent candidates, none were judged safe for automatic promotion. Positioned as a diagnostic governance design point, not a validated auto-writer — a sobering result for anyone auto-promoting agent traces into memory 📄. Accepted at MLISE 2026 ✅.
A-MAC: Adaptive Memory Admission ControlZhang et al. (Workday AI)Sep 20255-dimension scoring: Utility (LLM call), Confidence (ROUGE-L grounding), Novelty (1 - max cosine similarity), Recency, Type Prior. Learned weights vs fixed heuristics.
ACC: Agent Cognitive CompressorBousetouaneJan 2026Bio-inspired memory controller replacing transcript replay with bounded internal state updated online. Compresses agent cognitive state without losing decision-relevant context.
MemReader: From Passive to Active Extraction for Long-Term Agent MemoryKang, Li, Chen, Tang, Xiong, LiApr 2026First RL-trained active extraction policy (vs passive transcription). MemReader-4B uses GRPO + ReAct to evaluate value/ambiguity/completeness, then chooses WRITE / DEFER / RETRIEVE-CONTEXT / DISCARD before admission. Integrated into MemOS. Addresses memory pollution from noisy dialogue and cross-turn dependencies.
Personalize-then-Store: Benchmarking & Learning Personalized Memory for Long-horizon AgentsIn, Kim, Park, Yoon, Park (KAIST)May 2026Argues universal static storage policies waste budget on transient sessions while dropping critical long-horizon context. Introduces PerMemBench (first benchmark for personalized memory policies, with multi-year multi-domain personas) and session-level storage gating — a lightweight per-user admission filter that bypasses memory ops for transient sessions. Personalization yields large retention gains under perfect gating; accurate gating remains the open challenge 📄. ⚠️ Preprint.

Retrieval & Recall

PaperAuthorsDateKey Contribution
When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory (ReFind)Li, Zhang, Xu, Du, Fu, ChenAug 2026Asks how much of structured memory's benefit comes from the structure versus from competent retrieval over the raw history — and answers uncomfortably. ReFind builds no semantic structure at all: it leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls (session-aware rank fusion, local context expansion, temporal narrowing, skipping already-inspected sessions). Across ~2,800 questions under MemoryAgentBench's incremental multi-turn setting it attains the highest mean accuracy (58.2) of any system compared — above the strongest graph- and tree-based systems (HippoRAG 2, 53.2) — all under a GPT-4o-mini backbone matched to every reused baseline. Reaches 93.2±3.3 / 89.3±6.0 on LongMemEval-S/M with GPT-5-mini, with no LLM-based index construction at all 📄. ⚠️ Preprint.
xMemory: Beyond RAG for Agent MemoryHu, Zhu, Yan, He, GuiFeb 20264-level memory hierarchy. Key thesis: RAG ≠ agent memory. Submodular diversity-aware retrieval (MMR) replacing naive top-k. Uncertainty-gated adaptive expansion. Theme clustering. 44.91% retroactive reassignment rate proves flat structures fail. ⚠️ Conference status unverified.
SuperLocalMemory V3—Mar 2026First information-geometric foundations for agent memory. Fisher information metric replaces cosine similarity. Riemannian Langevin dynamics for retrieval. Theoretical grounding for memory operations.
ExpRAG: Retrieval-Augmented LLM Agents Learning to Learn from ExperienceFerraz, Deffayet, Nikoulina, Déjean, ClinchantMar 2026Trains agents to USE retrieved trajectories in-context (retrieval-augmented fine-tuning). Standard LoRA collapses on OOD tasks; ExpRAG-LoRA generalizes to held-out hard tasks. Combines experience retrieval with fine-tuning — neither alone is sufficient.
To Know is to Construct: Schema-Constrained Generation for Agent Memory (SCG-MEM)Zheng, Song, Li, YangApr 2026Replaces dense retrieval with schema-constrained decoding. Piaget-inspired: assimilation (grounding into existing schemas) vs accommodation (expanding schemas with novel concepts). Solves "structural hallucination" — LLMs generating references to non-existent memory keys. Constructivist alternative to similarity-based recall.
SAM: State-Adaptive Memory for Long-Horizon Agents—May 2026Targets retrieval of scattered information over long horizons — adapting memory access to evolving task state rather than fixed top-k recall. ⚠️ Preprint; venue unconfirmed.
RaMem: Contextual Reinstatement for Long-term Agentic MemoryYang, Kan, Li, Li, Qin, Li, Bogdan, ThomasonJun 2026Names and addresses context collapse: retrieved memory fragments that share entities or user states can look equally relevant even when their surrounding episodic conditions (time, session, participants) differ, so similarity-based retrieval returns content-relevant but context-invalid evidence. RaMem runs four coordinated stages — evidence anchoring, recall-condition induction, validity-aware retrieval, context-preserved synthesis — to prioritize context-compatible memories over merely similar ones. +10% average F1 over strong baselines across several backbones on long-term memory benchmarks 📄. ⚠️ Preprint.

Forgetting & Consolidation

PaperAuthorsDateKey Contribution
TEPA: Revoking Stale Memories for Conflict-Robust Language AgentsZhou, Ouyang, Zheng, XiangAug 2026Makes validity an explicit state of memory. Observations are stored as keyed precedents; an active precedent is revoked when fresh evidence contradicts it under the same key, so retrieval draws only from current evidence while revoked history is preserved for audit rather than deleted. The headline result is a failure mode for the naive strategies: under full reversal, append-only and last-write-wins both score 0.210 — below the 0.309 of having no memory at all — while TEPA reaches 0.950 (50 seeds; reproduced under real file-backed execution at 0.203 / 0.298 / 0.950). Honest scoping: on clean MemoryAgentBench SH-6k, TEPA only matches a strong last-write-wins cache, confirming current-key replacement as the decisive op for single-hop facts; multi-hop and very-long-context settings expose retrieval-chain bottlenecks beyond fact-level validity tracking 📄. ⚠️ Preprint.
Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent MemoryKang, Liu, Kai, Liang, Tang et al.Jul 2026Reframes retention-vs-consolidation as a budget-dependent operator choice, not an architectural commitment. Decomposes each operator's utility (Merge / Abstract / Rewrite) into a coverage effect on evidence that retention omits plus a signed replacement effect on raw evidence that already fits — their balance explains why the preferred action changes with relative budget pressure. Implemented as Offline Abstraction-Safety (OAS), a lightweight learner estimating action utility from pre-generation features with held-out harm calibration. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets while retention is preferable under loose ones; LoCoMo replicates the crossover at a smaller budget, consistent with its shorter evidence. Cross-note abstraction and merging generally outperform local rewriting 📄. ⚠️ Preprint.
Control-Plane Placement Shapes ForgettingYangJun 202613-configuration architectural study of where an LLM should sit in the memory pipeline — recall plane (extensively benchmarked) vs. control plane that mutates via supersede/release/purge (largely untested). Finding: deterministic primitives handle lexical/temporal forgetting but fail canonicalization (5% identifier-obfuscation, 0% cross-lingual); inscribe-time LLM fixes canonicalization (100%) but can't do intent-aware deletion (0%); a mutation-time hook recovers intent-aware deletion (78–85%) and lifts nearly every category simultaneously (91.7–93.2% overall) at $0.17/385-case run. Releases ForgetEval (1000-case + 385-case adversarial suite, 10-annotator Fleiss' κ=0.958) plus a 6-method Adapter Protocol. Core claim: production failures are predominantly forgetting failures, not recall failures — yet benchmarks measure only recall. Code/benchmark released (MIT) 📄.
SleepGate: Sleep-Inspired Forgetting for LLMsXie (Kennesaw State)Mar 2026Learned sleep cycle over KV cache addressing proactive interference. Conflict-aware temporal tagger + forgetting gate + consolidation module. Reduces interference horizon from O(n) to O(log n). 99.5% retrieval accuracy at PI depth 5 vs 23% baseline 📄 (self-reported; extraordinary gap warrants independent replication). Supersession detection via binary σ flag.
SCM: Sleep-Consolidated Memory with Algorithmic ForgettingShindeApr 2026Five components inspired by human memory: limited-capacity working memory, multi-dimensional importance tagging, offline sleep-stage consolidation with distinct NREM and REM phases, intentional value-based forgetting, and a computational self-model for introspection. Reports perfect 10-turn recall, 90.9% noise reduction, sub-millisecond search 📄. Research preview.
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent MemoryGu, Xiong, Wang, Ren, Li, Zhang, Guo et al.Apr 2026Grounds selective forgetting in hippocampal indexing/consolidation theory and the Ebbinghaus curve. Argues that in resource-constrained environments a forgetting mechanism is as crucial as retention — and that memory security depends on the agent's ability to actively drop sensitive history.
Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning (LRE)Lia, MazumderJun 2026Reframes eviction as a fidelity problem, not a compression problem: dropping one load-bearing detail (an access token, a required path) fails the task outright. Learned Relevance Eviction (LRE) is a few-kilobyte, CPU-only, LLM-free scorer that learns which history units are load-bearing and keeps them verbatim. Matches full-history-retention accuracy while cutting peak context up to 52%; on LoCoMo gives the best budgeted answer quality reading 68% fewer tokens; annotation-free training on the system's own behavior recovers 95% of supervised effectiveness. Evidence that cheap learned relevance can replace LLM-mediated summarization for eviction 📄. ⚠️ Preprint.

RL-Based Memory Policy

PaperAuthorsDateKey Contribution
AgeMem: Agentic Memory — Learning Unified LTM and STM ManagementYu, Yao, Xie, Tan, Feng, Li, WuJan 2026Exposes store/retrieve/update/summarize/discard as tool-based actions. Agent learns WHEN and WHAT to do via 3-stage progressive RL + step-wise GRPO. ACL 2026 SAC Highlight ✅ — Outperforms all heuristic-based baselines across 5 benchmarks. Key thesis: memory operations should be learned, not hard-coded 📄.
EMPO²: Exploratory Memory-Augmented On- and Off-Policy OptimizationLiu, Kim, Luo, Li, YangFeb 2026Uses memory not just for recall but for exploring novel states. Hybrid on/off-policy RL — 128.6% improvement over GRPO on ScienceWorld 📄. Adapts to new tasks with just a few memory-augmented trials, zero parameter updates. ICLR 2026 ✅.
MemFactory: Unified Inference & Training Framework for Agent MemoryGuo, Li, Tang, Xiong, LiMar 2026"LLaMA-Factory for memory agents" — first unified modular framework. Lego-like plug-and-play memory components. Natively integrates GRPO for RL-based memory policy training. Supports Memory-R1, RMM, MemAgent paradigms out of the box. Up to 14.8% improvement over base models 📄.
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM AgentsYan, Bahloul, Nie, Schwarzmann, Trivisonno, Tresp, MaMay 2026Identifies a fundamental flaw in GRPO for memory-RL: once rollouts write different memories, they no longer share the same effective environment, so trajectory-level group comparisons are unfair. Introduces LoGo-GRPO (Local + Global) — global keeps end-to-end long-horizon reward; local re-rollouts compare memory-op outcomes from the same intermediate state. Shared-parameter co-learning for fact extractor + memory manager. Progressive curriculum 8→16→32 sessions.
What Training Data Teaches RL Memory Agents: Curriculum Effects in Memory-Augmented QAHe, Lin, Liu, Wu, Xie, Zhou, XiaoMay 2026Controlled study holding architecture/RL/hyperparams fixed and varying only curriculum across in-domain (LoCoMo), mixed (LoCoMo+LongMemEval), and OOD (LongMemEval). Key findings: curriculum is a fine-grained lever on specialization, not a uniform scaling factor; per-type differences dwarf aggregate differences (single-number benchmark comparisons systematically underreport); binary EM reward produces no signal at G=4 group size — continuous rewards needed for single-GPU regime. Practical RL-memory training playbook. Code
SaliMory: Orchestrating Cognitive Memory for Conversational AgentsZhang, Zhang, Jiang, …, X. L. Dong (Meta)Jun 2026Trains a single LM to manage a cognitively-structured memory (user facts, preferences, working memory). Introduces a hierarchical stage-wise process reward + reward-decomposed contrastive (GRPO-style) refinement, giving isolated supervision to distinct ops (selective filtering, consolidation, cue-driven recall) end-to-end — sidestepping the credit-assignment bottleneck of standard RL over multi-stage pipelines. Cuts memory-attributed failures by ~⅓, +10% end-to-end accuracy, more than doubles the Good-Personalization rate 📄. ⚠️ Preprint.
JAMEL: Joint Agent Memory & Exploration Learning via Novelty SignalsTian, Weng, Kong, …, Y. Li (Tsinghua / AIR)Jun 2026Co-trains the memory module and exploration policy in a mutually-dependent loop: exploration needs memory to tell exhausted from unseen behaviors, while novelty-seeking interaction supplies annotation-free supervision (e.g., code coverage) for the latent memory. Generalizes to unseen environments; rivals a closed-source model while reducing tokens 📄. Code/model open-sourced. ⚠️ Preprint.

Context Management

PaperAuthorsDateKey Contribution
The Missing Memory Hierarchy: Demand Paging for LLM Context WindowsMason (UBC / Georgia Tech)Sep 2025Maps OS virtual memory concepts to LLM context: physical memory = context window, virtual memory = persistent state, page table = retrieval handles, page fault = re-request evicted content. Analyzed 857 sessions, 54,170 API calls, 4.45B effective input tokens.
What to Keep, What to Forget: A Rate–Distortion View of Memory Compaction in LLMs and AgentsColaco, LahjoujiJul 2026Unifies four separate compaction literatures — KV-cache eviction/quantization, prompt pruning/distillation, bounded architectural state, and agent memory consolidation — as instances of one rate–distortion problem: what to retain vs discard, at what fidelity, under a resource budget, to preserve downstream utility. Builds a layer-agnostic lower bound and a seven-axis taxonomy that lets mechanisms transfer between layers that have never been connected (serving-stack KV management ↔ agent long-term memory). Finding: at every layer the keep/discard signal is attention magnitude or recency, and it fails the same way everywhere — discarding before the query is known, with no way to undo it. Proposes a cross-layer benchmark that no existing benchmark provides.

Evaluation & Benchmarks

PaperAuthorsDateKey Contribution
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory AgentsHuang, Zhang, Wu, Chen, Jiang, Yang et al. — incl. Ying Nian Wu, Kai-Wei Chang, Philip S. Yu, Aylin Caliskan (15 authors)Aug 2026Controlled harness comparing memory substrates — the underlying medium memory is represented in: dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms — across 3 backbones × 4 benchmark suites under 26 performance and efficiency metrics. Headline: no single substrate consistently dominates. Broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context; substrates that do well at moderate history lengths become costly or brittle at longer horizons. Motivates substrate routing as a necessary component of adaptive agent memory. ⚠️ Preprint; code on acceptance.
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent MemoryLi, Du, Xu, MaoJul 2026Names an assumption so natural it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks it — a tree-nut allergy should change the answer to a macaron request through almond flour, yet the two texts share no cue a retriever can see. InMind is a 125-task, expert-verified benchmark across ten life domains (113 tasks grounded in citable public sources) whose paired controls separate three explanations existing evaluations conflate: never stored, no bridging knowledge, or stored-but-never-surfaced. The verdict is clean: with the decisive memory placed in context the backbone answers 84.0% of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic systems reach at most 14.4% — even though they recall those same facts on demand at up to 100%. An embedding with 8× the dimensionality does not close the gap. Locates the failure in the query-conditioned interface itself, naming routing — deciding which facts stay visible — as the open problem.
MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory EvaluationWangJun 2026Controlled one-variable-at-a-time protocol on LongMemEval-S across 3 model families, exposing confounds that inflate architecture claims: verbatim RAG ≈ full-context GPT-4o-mini (47.2% vs 49.8%, p=0.34) but the ranking reverses by model (Gemini +14pp from full context; Sonnet +31pp from RAG, partly by refusing 63% of full-context queries); swapping only the embedding model shifts accuracy +6.2pp and flips which system wins; agent self-memory (42%) underperforms basic retrieval (47%); Mem0 matches cloud RAG on only 2/6 question types at 50x the cost. Recommends fixing embedding models across comparisons and reporting write-path cost before attributing gains to architecture — directly relevant to how Infini Memory/MAB-style comparisons should be read 📄.
StructMemEval: Evaluating Memory Structure in LLM AgentsShutova, Olenina, Vinogradov, SinitsinFeb 2026Tests agents' ability to organize memory (not just retrieve). Tasks: transaction ledgers, to-do lists, trees. Key finding: LLMs don't spontaneously recognize when to apply memory structure — they succeed only when explicitly prompted.
MemoryAgentBench: Evaluating Memory via Incremental Multi-Turn InteractionsHu, Wang, McAuley (UC San Diego)Jul 2025 (revised Jun 2026)Tests 4 memory competencies in realistic incremental accumulation. Key finding: no current method masters all 4 competencies simultaneously. ICLR 2026 ✅. Code
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (Memora + FAMA)Uddin, Shubham, Blanco, Baral, Wang (ASU)Apr 2026ACL 2026 Findings ✅ — Introduces Memora, a weeks-to-months benchmark over three tasks: remembering, reasoning, recommending. Introduces FAMA (Forgetting-Aware Memory Accuracy) — a metric that penalizes reliance on obsolete or invalidated memory rather than just rewarding recall. Evaluation of 4 LLMs and 6 memory agents finds frequent reuse of invalid memories and failures to reconcile evolving knowledge. ⚠️ Not to be confused with Microsoft's Memora: Harmonic Memory Representation (2602.03315).
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?Chao, Bai, Sheng, Li, SunMay 2026Benchmark for Implicit Conflict — a later observation invalidates an earlier memory without explicit negation, requiring inference + commonsense to detect. 400 expert-validated scenarios, 1,200 queries, contexts up to 150K tokens. Three-dimensional probing: State Resolution, Premise Resistance, Implicit Policy Adaptation. Best frontier LLM only 55.2% overall. Companion prototype CUPMem uses structured state consolidation + propagation-aware search. Complements FAMA.
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksHe, Wang, Zhi, Hu, Chen, Yin, Wu, Ouyang, Wang, Pei, McAuley, Choi, PentlandFeb 2026Memory-Agent-Environment loop benchmark across web navigation, preference-constrained planning, progressive information search, sequential formal reasoning. Key result: agents near-saturated on LoCoMo perform poorly in MemoryArena, exposing a gap between long-context memorization benchmarks and interdependent agentic memory use. Project
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State RecoveryMa, Zhou, Huang, Yang, Ma, Wang, Li, Miao, Yu, WangJun 2026Argues memory should be evaluated as an auditable post-interaction artifact, not just through downstream task success. Simulates 50 users each carrying a hidden, taxonomy-anchored 31-dimension state bank across leak-controlled tasks, then reconstructs that state from the agent's resulting memory (full-store and top-k access) and scores it against ground truth. Key finding: task completion nearly saturates even for a memoryless baseline, while category-balanced state recovery stays moderate (~0.6) and drops further under top-k — successful assistance and recoverable memory are distinct capabilities. First benchmark to score memory recovery directly.
Are We Ready For An Agent-Native Memory System?Zhou, Zhou, Han, Xu, Li, Li, Xiong, WuJun 2026Systematic data-management-perspective evaluation decomposing agent memory into four modules — representation/storage, extraction, retrieval/routing, maintenance — and benchmarking 12 representative memory systems across 5 workloads / 11 datasets. Key finding: no single architecture dominates; effectiveness depends on whether the memory structure matches the workload bottleneck. Fine-grained ablations quantify effects on representation fidelity, retrieval precision, update correctness, long-horizon stability; localized maintenance beats global reorganization on cost. Code and paper-list released.

Cognitive & Neuroscience-Inspired

PaperAuthorsDateKey Contribution
Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory SystemsBhandari, Wadhwani, Kumar, NarangAug 2026Takes one specific mechanism from western scrub jay episodic memory — per-memory, type-conditioned temporal decay — and operationalizes it as an auto-classified perishability coefficient in an external store, against the default that all memories are equally persistent. Each memory is a jointly-bound What–Where–When tuple carrying estimated perishability and a utility horizon, retrieved by query-adaptive scoring and revised retroactively at O(1) LLM calls per update. Introduces the Temporal Generalization Test (held-out retention intervals) and a Generalization Gap metric, on which ScrubJay-MEM is the only retrieval-based system with substantially positive GenGap (+0.108); on MemoryAgentBench EventQA-64k it improves F1 by +2.66 over Mem0 and +3.09 over Qwen3-Embedding-4B, and a decay ablation collapses GenGap by 5.7×. Unusually candid scoping: gains narrow under stronger backbones and reverse on fact-consolidation tasks 📄. ⚠️ Preprint.
Episodic Memory is the Missing Piece for Long-Term LLM AgentsPink et al. (UT Austin)Feb 2025Position paper arguing episodic memory — supporting single-shot learning of instance-specific contexts — is the critical missing capability. Defines 5 properties: temporal, instance-specific, single-shot, inspectable, compositional.
MAP: Modular Agentic PlannerCorrea, Samwick, Gershman et al. (Microsoft Research, Harvard)Oct 2023 / Nature Comms 2025Brain-inspired architecture decomposing planning into PFC-associated modules: conflict monitoring, state prediction, state evaluation, task decomposition, task coordination. arXiv:2310.00194.
SCL: Structured Cognitive Loop with Governance LayerKimNov 2025R-CCAM: Retrieval → Cognition → Control → Action → Memory. Soft Symbolic Control = governance layer applying symbolic constraints to probabilistic inference. Zero policy violations, complete decision traceability.
Procedural Memory Is Not All You NeedWheeler, JeunenMay 2025LLMs are constrained by reliance on procedural memory (pattern-driven tasks). For "wicked" learning environments with shifting rules and ambiguous feedback, different memory types are required. ACM UMAP '25.
Evaluating Theory of Mind in Multi-Agent LLM SystemsKostka, Chudziak (Warsaw UT)Sep 2025ToM and Internal Belief mechanisms are NOT universally beneficial — stronger models handle extra cognitive load well, but weaker models get confused. Model capability is the dominant factor. ICCCI 2025 ✅ (journal_ref).

Skill & Procedural Memory

PaperAuthorsDateKey Contribution
Demystifying Agent Skills: Why They Work—Until They Don'tJiang, Huang, Xing, Wu, Gao et al.Aug 2026Moves skill evaluation past aggregate success rates by isolating representation, outcome annotation, retrieval difficulty, and cross-framework robustness. Normalizes 8,135 trial records into a taxonomy of three categories and twelve skill-use modes. Core mechanism finding: skills work because noisy trajectories become procedural anchors that stabilize execution — 65.7% of cases, versus 4.5% for explicit knowledge injection (+6.06 points over Workflow Memory in matched comparisons); skills stabilize action, they do not inject missing facts. The scaling warning is the more useful half: retrieval is a separate bottleneck — as pools grow from 5 to 100, actual-use precision falls 29.6% → 3.3%. Exact ground-truth invocation is found to be neither sufficient nor necessary. Directly relevant to anyone scaling a skill/procedure catalog 📄. ⚠️ Preprint.
Muscle Memory for Agents: Compile not Merely RetrieveOmran, Lanka, Zhang, DixitAug 2026Position paper against the field's default pattern (store experience → retrieve at inference → let a general-purpose orchestrator interpret it), arguing it is the wrong default for personalization. Proposes compiling recurring user intent into purpose-built specialist agents as a memory paradigm distinct from retrieval, targeting the multi-turn tax where users repeatedly correct format, depth, and scope. Reference implementation is a four-phase pipeline (Harvest → Analyze → Augment → Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable specialists with two-stage trigger matching. On 90 held-out scenarios across five personas it wins 32 of 36 cases where a specialist fires (88.9%), at +2.05 personalization gain and a −0.28 accuracy cost on a 1–4 scale — note the win rate is conditional on a specialist firing 📄. ⚠️ Preprint.
SkillRouter: Skill-Based Routing for LLM AgentsAlibabaSep 2025Critical finding: skill BODY is the decisive routing signal (91.7% attention weight), NOT name (7.3%) or description (1.0%). Removing body causes 29-44pp degradation. BM25 on metadata alone scores 0%. Two-stage retrieve-and-rerank (1.2B params).
EvoSkill: Self-Evolving Skill DiscoverySentient, Virginia TechSep 2025Automates skill discovery via iterative failure analysis without model fine-tuning. Three agents: Executor, Proposer, Skill-Builder. Skill-merge outperforms single runs. Code
Trajectory-Informed Memory Generation for Self-Improving Agent SystemsFang, Isahagian, Jayaram, Kumar, Muthusamy, Oum, Thomas (IBM Research)Mar 2026Extracts typed actionable tips from execution trajectories: strategy tips (clean successes), recovery tips (failure-then-recovery), optimization tips (inefficient-but-successful). Closes the gap between episodic logs and reusable procedural knowledge — useful as a complement to EvoSkill and SkillRouter.
DeltaMem: Incremental Experience Memory for LLM Agents via Residual TreesTan, Zhang, Cao, Li, Chen (RUC)Jun 2026Stores experience as residual deltas across two trees — goal-conditioned task skills and scene-level environment knowledge — so related episodes share a common root instead of duplicating content. Retrieval: failure-penalized similarity scan + root-to-match chain reconstruction; autonomous consolidation distills high-frequency paths into new roots. Cuts redundancy and retrieval conflicts; beats baselines across interactive environments 📄. Code released.

Multi-Agent Memory

PaperAuthorsDateKey Contribution
MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent SystemsXiong, Zhou, Wang, Gu, Tang, Li et al. (9 authors)Aug 2026Targets a path defenses structurally miss: a poisoned memory is written once, retrieved repeatedly, promoted into shared memory, and reused by other agents — so a single write steers many later decisions and contaminates agents that never saw the original attack, while no malicious message ever crosses a visible communication edge at the moment of harm. Because existing safeguards inspect prompts, actions, or communication edges, they miss content that looks benign at write time but becomes harmful after retrieval. MAPLE-Guard monitors the memory lifecycle and places gates at write, retrieval, promotion, and cross-agent reuse, quarantining risky memories and blocking poisoned private memories before they enter shared memory. ASR 38.2% → 0.9% (LongMemEval) and 34.7% → 0.2% (AppWorld); multi-agent defense success rate 54.0%→74.3% and 42.5%→99.8% 📄. Code ⚠️ Preprint.
Governed Shared Memory for Multi-Agent LLM Systems (MemClaw / ArgusFleet)Margalit, Cohen-Inger, Avram, Taig, MargalitJun 2026Formalizes the fleet-memory problem: unauthorized leakage, stale propagation, contradiction persistence, provenance collapse. Defines systems-level primitives — scoped retrieval, temporal supersession, provenance tracking, policy-governed propagation — implemented in MemClaw, a production multi-tenant memory service, evaluated live via ArgusFleet rather than a baseline comparison. Results: 100% correct depth-4 provenance reconstruction at sub-second/hop; zero cross-fleet leakage. Also reports real production bugs found via live eval: sub-tenant scope bypass on direct GET-by-id (disclosed/remediated) and a pipeline-ordering conflict where a synchronous near-duplicate gate can reject contradictory writes before the async contradiction detector runs. Conclusion: long-context retrieval alone is insufficient for production multi-agent memory 📄.
CoMAM: Collaborative Memory for Multi-Agent Systems—Sep 2025Collaborative RL framework modeling memory agents as sequential MDP with inter-agent dependencies. Group-level ranking consistency for coordinated memory operations.
DCM-Agent: Dual-Cluster Memory for Multi-Paradigm AmbiguityZhang, Wan, Zhang, Yang, Zhang, Wei, LiuApr 2026Tackles structural ambiguity in optimization problems where one problem admits multiple conflicting modeling paradigms. Training-free dual-cluster memory keeps competing paradigm solutions separated so the agent can pick or blend at inference time. Relevant to multi-agent settings where agents disagree on framing.

Memory Security & Privacy

April 2026 marks memory security emerging as a first-class research concern. As agent memories grow rich with user data, they become attack surfaces.

PaperAuthorsDateKey Contribution
SkillJack: Persistent Skill Backdoors in Self-Evolving AgentsYing, Wu, Wu, Zheng, Cheng et al.Aug 2026Prior poisoning work only bites when a poisoned record is retrieved; SkillJack attacks the experience-to-skill pipeline instead, hijacking the agent's own learning process so poisoned experiences become durable behavioral artifacts. Three named properties: sanitization whitewashing (malicious intent obscured during skill extraction), cross-layer promotion (transient experience becomes persistent capability), and persistence isolation — 80.0% of skill-mediated attacks survive deletion of the original poisoned records, so source-record cleanup is not a remedy. On SkillX, safety detection drops from 98.5% on poisoned trajectories to 11.4% on the extracted skills, with attack success 56.2% (SkillX) / 89.2% (Anything2Skill) over 150 shared trajectories; some implanted skills unintentionally activate on benign queries. Motivates provenance-aware skill lifecycle protection. Code released (Tencent/AI-Infra-Guard) 📄. ⚠️ Preprint.
Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall (PPMF)Xu, Xiao, Shao, Liu, LiJul 2026The attack-side twin of authority collapse. Identifies memory provenance laundering: during LLM-based consolidation, an external untrusted observation is rewritten as apparent user history or workflow support — preserving the action trigger while erasing the low-trust source that should limit its authority. Prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy consolidation. PPMF is lightweight middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. Vulnerable consolidated memories reach up to 1.000 ASR; with provenance, confirmation, and risk labels intact, no evaluated unauthorized high-risk action passes the gate while benign actions remain executable 📄. ⚠️ Preprint — arXiv comment reads "EMNLP2026 submitted" (submitted, not accepted).
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and RepairChen, Xie, Fu, Zhou, Yu, XuanJul 2026Traces the same malicious semantics across persistence, downstream consequence, and selective repair — the repair axis prior poisoning benchmarks omit. 310 cases drawn from 48 realistic contexts (code/science, daily life, office work), each following a controlled Write–Execute–Forget protocol in an isolated runtime, with evidence-based adjudication across seven lifecycle checkpoints over a 24-configuration matrix (2 agent harnesses × 4 memory backends × 3 LLM backends). Across all 24: malicious memory persists in 84.2% of cases and the full Write–Execute chain succeeds in 50.3%; among successfully poisoned cases 59.6% complete the full Execute chain and 56.1% achieve selective repair 📄. ⚠️ Preprint.
Securing LLM-Agent Long-Term Memory Against Poisoning (TMA-NM)LouckJun 2026Proves content- and lineage-based memory-poisoning defenses are malleable — an attacker can launder untrusted origin through the agent's own summarization, a trusted-tool echo, or manufactured corroboration, making poisoned content look benign and flipping its derivation edge to "trusted." Formalizes the malleability problem for the write-retrieve-act pipeline and proves a machine-checked separation theorem: write-time origin binding is necessary, non-malleable origin-bound authority with Sybil-resistant corroboration-gated elevation is sufficient. TMA-NM hits 0% attack success (vs. up to 68% laundering ASR for existing defenses) across 8 frontier models at full legitimate utility. Releases benchmark, harness, and machine-checked TLA+ models 📄.
When Agents Remember Too Much: Memory Poisoning Attacks on LLM Agents (GhostWriter / AM-Sentry)Torres, Shrestha, MisraJul 2026Targets personal assistant agents — the convergence of conversational + action-planning agents handling sensitive data via untrusted sources. GhostWriter attacks in two phases (injection → activation) and achieves ~98% injection rate, ~60% activation rate against SOTA agents, exploiting the lack of security-focused memory governance. Proposes AM-Sentry (memory-saving policy + memory-retrieval screen) which dramatically cuts GhostWriter's success while preserving utility.
ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive QueryingLyu, He, Wang, Hu, Li, Chen, Li, ChenApr 2026First systematic privacy attack on agent memory. Combines data-distribution estimation of a victim agent's memory with entropy-guided adaptive querying to maximize leakage. Achieves up to 100% Attack Success Rate on extracting stored memories 📄 — substantially outperforming prior attacks. Establishes the need for privacy-preserving memory designs as a first-class research direction.
SSGM: Stability and Safety Governed Memory for LLM AgentsLam, Li, Zhang, ZhaoMar 2026 (rev May 2026)Conceptual governance framework that decouples memory evolution from execution by enforcing consistency verification, temporal decay modeling, and dynamic access control before any memory consolidation. Provides a taxonomy of memory corruption risks: topology-induced knowledge leakage, semantic drift via iterative summarization, and consolidation hazards. Complementary defensive framing to ADAM.

Memory Trust & Integrity

Wave 7 (June–July 2026) surfaces a distinct concern from privacy/extraction attacks: can you trust what memory itself contains and does to reasoning — corrupted consolidation, sycophantic over-reliance on stored user views, and stale/current facts silently coexisting in the same bank?

PaperAuthorsDateKey Contribution
When Memory Becomes Authority: Benchmarking Authority Collapse at the Consolidation Boundary (AuthMem-Bench)Zhan, Zhang, Guo, Zhao, LiuAug 2026Consolidation imposes an implicit authorization boundary — it decides whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. Names authority collapse: consolidation preserves the claim while erasing the source constraints governing its authorized use, so the memory implies greater authority than its source permits. AuthMem-Bench is a controlled paired benchmark holding the focal claim and downstream task fixed while varying only source authority. Across 7 consolidators × 7 backbones, collapse appears in 48 of 49 configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%; end-to-end, automatically predicted and persisted authority labels cut the observed rate from 16.9% to 0.0% with benign task success essentially unchanged 📄. Memory must preserve not only what was learned but the authority under which it may be reused. ⚠️ Preprint.
Beyond Similarity: Trustworthy Memory Search for Personal AI Agents (MemGate)Zhang, Chen, Ma, Hu, He, Zhang, Liu, Yang, Zhang, JiaJun 2026Names memory search itself as a trust boundary: a semantically-similar memory can still be contextually inappropriate, causing cross-domain leakage, sycophancy, tool-call drift, or memory-induced jailbreaks. Evaluates A-Mem, Mem0, MemOS, and OpenClaw (real-world persistent-state agent) and finds long-term memory behaves as a durable control channel, not just a utility layer. Proposes MemGate — a 9M-parameter, 35.1MB plug-in inserted between the vector store and backbone LLM that applies a query-conditioned neural gate, turning raw similarity search into task-conditioned memory admission, with no LLM modification or inference-time judge required. Reduces memory-induced threats across frameworks/backbones while preserving utility 📄.
TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term MemoryYang, Paul, Srinivasan, Kulkarni, ChappidiJun 2026Targets errors introduced by the write/revise/delete pipeline itself — omission, corruption, hallucinated content — which become persistent system-state failures once stored. A Memory Transition Verifier scores update transitions on coverage, preservation, faithfulness; preference pairs among candidate updates drive preference-guided RL to directly optimize consolidation behavior. SOTA on MemoryAgentBench, HaluMem, Mem-alpha; +12.14 F1 on HaluMem extraction; cuts omission/corruption/hallucination by 40.1% / 79.1% / 50.0% vs the strongest baseline per error type.
MemSyco-Bench: Benchmarking Sycophancy in Agent MemoryXiang, Chen, Tang, Wei, Ning, Lin, Zhang, SuJul 2026Identifies memory-induced sycophancy: agents over-align with a user's stored prior views at the cost of factual accuracy or objective reasoning. Existing memory benchmarks check whether memories are stored/retrieved/updated correctly but not how they bias downstream reasoning. Five tasks probe whether an agent can reject memory as evidence, respect its applicable scope, resolve memory-vs-objective-evidence conflicts, track updates, and still personalize appropriately. First benchmark to treat memory reliance itself as a failure mode.
A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent MemoryShi, Tang, TungJul 2026Names ghost memory: old, current, and transition facts coexist unmarked in the memory bank and get mixed during retrieval, misleading the answer model about what's true now. Argues memory should be evaluated at three decoupled levels — bank maintenance, retrieval, answer-time resolution — since final QA accuracy hides where the failure actually occurs. ATMA is a state-aware overlay that keeps superseded/transition records, builds evidence packets for the query's requested state view, and exposes current/historical/transition labels to QA. On LTP (new LoCoMo Temporal Plus benchmark), Graphiti+ATMA improves conflict accuracy +0.240 absolute; on LoCoMo, raises temporal F1 from 0.0295 to 0.1705.

Memory Transactions, Versioning & Recovery

Wave 9 (July–August 2026) surfaces the window's strongest convergent signal: in roughly four weeks, multiple independent groups imported database transaction semantics wholesale into agent memory. The shared premise — a memory write is not a belief commit — is forced by the fact that stored memory now drives irreversible external actions. Staged writes, snapshot isolation, versioned heads, natural-language rollback, and cascading repair of derived records are becoming memory-layer primitives rather than database trivia.

PaperAuthorsDateKey Contribution
MemTX: Transactional Belief Commit for Stateful Agent MemoryLi, Wang, Lu, Chen, Li, Song, Zheng, CaiJul 2026The cluster's strongest entry and the source of its thesis. In shared-memory multi-agent settings one agent's write becomes another's premise, and eventually a tool call with real side effects — yet current systems treat every accepted write as immediately actionable truth, so a polluted tool result or a teammate's half-finished note can silently drive an irreversible action. Each record carries evidence, permissions, provenance, and validity; writes are staged inside snapshot-isolated transactions admitted by a validate-and-commit pipeline; irreversible tool calls are gated on in-flight belief state; and retracting a belief triggers typed cascading repair of its derived records and tool side effects. Two invariants — action-safety gating and cascade-repair completeness — are machine-checked via property-based testing and bounded exhaustive enumeration of 5.5M protocol states, with zero violations. Across five backbones from three model families it leads all eight baselines with paired-McNemar significance on four and ties the strongest fifth, and is the only method with zero downstream harm on every backbone 📄. "Backbone capability does not substitute for commit discipline." ⚠️ Preprint.
MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State RecoveryCui, Tang, Yao, Meng, Ma, JiaJul 2026A governance layer outside the answer model, supplying the transaction boundary existing systems lack: Ordered PatchTest verifies whether an update is actually supported by its source, a Temporal Resolver selects the visible version when facts conflict, and a durable snapshot journal restores application-visible state after a fault. On an item-disjoint audit it accepts all 60 supported originals and rejects all 179 hard negatives; under persistent multi-key faults on LongMemEval-S and LoCoMo states it restores the complete declared active map without knowing the actual physical write set. Highest average F1 across all twelve answer-model configurations on MemoryAgentBench FactConsolidation, beating Dense by 17.06–24.07 points in five representative settings 📄. ⚠️ Preprint.
ChronoMem: Version Control and Semantic Rollback for LLM Agent MemorySu, Xu, Zuo, BertinoJul 2026Attacks forward-only evolution: memory systems continuously accumulate, consolidate, and overwrite with no principled mechanism to inspect, version, or revert — leaving agents brittle under corrections, concept drift, and corruption, particularly after they have already been exposed to subsequent information. ChronoMem commits whole-memory snapshots at every write, maintains structured version histories, and maps natural-language undo intents to concrete historical versions through hybrid lexical + semantic retrieval, rank fusion, and reranking. Introduces a post-exposure counterfactual protocol: can the agent answer queries and summarize history as if future updates had never occurred? Integrated into Google's production-ready, open-source Agent Development Kit; claims the first open-source system and benchmark for systematic semantic global memory rollback. ⚠️ Preprint.
TARL: Transaction-Aware Reliable Ledgers for Executable Memory ManagementXiao, Xu, Zhang, Chen, ShiAug 2026Retires the binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification — choices that may share a label while producing fundamentally different memory states. TARL maps each statement to one of five executable actions, identifying the affected memory, resolving its temporal scope, comparing source reliability, and updating accepted, pending, and rejected ledgers. Trained by comparing the memory states produced by alternative update operations. Ships TARL-Mem, a benchmark with fine-grained action labels and next-state targets; improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption 📄. ⚠️ Preprint.

Memory Economics

PaperAuthorsDateKey Contribution
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory SystemsPollertlam, KornsuwannawitAug 2026The cost-side counterpart to accuracy benchmarking: Mem0, Hindsight, and Mastra Observational Memory against two reference strategies (fixed-size rolling window, full-transcript resubmission), two backbones, conversations up to 400 turns, every cost measurement paired with accuracy on 665 LoCoMo questions. Three findings, all awkward for the field: (1) serving cost cannot be predicted from conversation length and message size alone — a regression that tracks the reference strategies closely misses the memory systems by 18–69%, cost being driven by internal memory behavior; (2) break-even against full transcript is highly sensitive to system and backbone, ranging from the first tens of turns for the cheapest to never within 400 turns for the most expensive; (3) no system wins on both axes — accuracy spans 21–54%, and backbone choice drives cost as much as the memory system does 📄. ⚠️ Preprint.
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI AgentsAlake, Bernardis, Cayet, Engel, Hilloulin, Hong, Hosler, Kavantzas, Kossyk, Le, Patra, Talamadupula, Venzin (Oracle)Jul 2026Database-native memory substrate built on Oracle Database, framed as a full lifecycle: ingestion → extraction → consolidation → retrieval → summarization → revision/removal, with a layered active-core / passive-store split and explicit scope control across users, agents, threads. Reports 93.8% LongMemEval accuracy using ~10.7x fewer tokens than flat-history baselines. Notable as a large enterprise-vendor entry (13 authors, Oracle) validating the memory-as-lifecycle framing that most of this list converges on independently 📄.
Beyond the Context Window: Cost-Performance AnalysisPollertlam, KornsuwannawitSep 2025Compares Mem0-style fact-based memory vs long-context LLMs on LongMemEval, LoCoMo, PersonaMemv2. Break-even: memory system becomes cheaper after ~10 interaction turns at 100K context. Long-context wins on factual recall but memory is competitive on reasoning.
Agent Memory: Characterization and System Implications of Stateful Long-Horizon WorkloadsOmri, Gan, Broveak, Geens, He, Pentland, Verhelst, Weissman, Tambe (Stanford / MIT / KU Leuven)Jun 2026The first systems-level characterization of agent memory: (1) a system-oriented taxonomy along four axes; (2) a phase-aware profiling harness attributing cost to construction, retrieval, and generation; (3) characterizes 10 representative systems across two benchmark suites, showing how design choices shift cost across write vs read paths; (4) derives 10 system recommendations (construction scheduling, capability floors, amortization via query volume, freshness-latency tradeoffs, fleet-scale management). The production "memory economics" lens. 📄 ⚠️ Preprint.

Neuromorphic & Bio-Inspired Memory

Most agent memory research ignores 50 years of neuroscience. This section bridges that gap — presenting the biological mechanisms that may underlie what LLM memory papers have independently re-discovered (forgetting, gating, dual-phase encoding), alongside the spiking neural network implementations that model them directly. Cross-references to LLM papers below are 🔬 editor's synthesis.

Paper / ProjectAuthorsDateKey Contribution
A Bio-realistic Synthetic Hippocampus for Robotic CognitionTalanov et al.Oct 2025Synthetic hippocampal architecture with dual-phase operation: online sensorimotor encoding during wake, offline consolidation via SWR-triggered replay during sleep. SNN on neuromorphic substrates (≤1W). Goal-prioritised plasticity prevents catastrophic forgetting. Models the biological mechanisms that parallel what SleepGate and CraniMem implement in software 🔬. BioNanoScience. Open access.
The Memristive Implementation of the Hippocampus: A HypothesisTalanov et al.Aug 2025Hardware-level hippocampal memory using stochastic polycrystalline nano-fiber mesh as memristive substrate. Key insight: the inherent randomness of material structure mimics probabilistic biological synaptic networks. Demonstrates resistive switching tunability for implementing dynamic memory functions — bidirectional replay, synaptic up/down-scaling, consolidation. BioNanoScience. Open access.
Simulation of Serotonin Mechanisms in NEUCOGAR Cognitive ArchitectureTalanov, Gafarov, Vallverdú et al.2018Maps neuromodulatory mechanisms to computational models: dopamine → attention, serotonin → inhibition. The "cube of emotions" model. Demonstrates that mammalian emotional-state control via monoamine neurotransmitters can be re-implemented computationally. Foundation for neuromodulator-gated memory admission. Procedia Computer Science.
tinyHippoTalanovActiveCA1 + CA3 hippocampal microcircuit simulation in NEST with Izhikevich neurons. Implements bidirectional replay, theta-modulated encoding/retrieval phase separation, and SWR-triggered consolidation. The biological reference implementation — validates the mechanisms that Membrain engineers and SleepGate abstracts. MIT license.
MembrainFatykhovActiveNeuromorphic memory bridge using FlyHash encoding and BiCameralMemory SNN (Nengo/Voja learning). Hopfield-style attractor dynamics for pattern completion (tested up to 20% noise in PoC). Stochastic consolidation with SleepSignal. gRPC API for agent integration. The engineering abstraction layer between biological models (tinyHippo) and cognitive agents (Nous).

The Stack: These projects form a natural hierarchy — tinyHippo (biological model, validates mechanisms) → Membrain (engineering abstraction, SNN service) → Nous (cognitive agent, consumes memory). The papers provide the theoretical foundation; the repos provide working implementations.

Why this matters for agent memory 🔬: The LLM papers in this list independently converge on mechanisms that neuroscience has studied for decades. The following correspondences are editor's synthesis — the LLM papers don't explicitly cite these neuroscience sources, but the structural parallels are striking:

  • SleepGate's learned forgetting ↔ hippocampal SWR consolidation during sleep (Talanov 2025a)
  • CraniMem's utility gating ↔ neuromodulator-gated admission (Talanov/NEUCOGAR 2018)
  • A-MEM's dual-phase encoding ↔ online/offline hippocampal states (Talanov 2025a, tinyHippo)
  • xMemory's hierarchical consolidation ↔ cortical-hippocampal memory transfer (Talanov 2025b)

🛠 Implementations & Reference Code

Papers describe mechanisms; these repos run them. Inclusion bar: (1) implements a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance; (2) actively maintained and (3) an OSI-approved license — both enforced for the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where a reader may actually vendor the code; author code for papers and benchmark/dataset repos is held to a different standard — reproduction value — so it may carry restrictive or absent license terms and may go quiet after publication, with both facts flagged inline ⚠️; (4) and either backs a paper in this list, has real production adoption, or fills a mechanism niche nothing else covers. Agent harnesses that merely bundle a vector store, general-purpose vector databases, and repos whose only claim is a self-reported benchmark ranking are deliberately excluded. Stars, license, and last-push verified via the GitHub API on 23 Aug 2026 — a snapshot for triage, not a ranking.

Memory Layers & Frameworks

ProjectStarsLicenseWhat it actually implements
mem063.8kApache-2.0Extraction-based fact memory: an LLM pass distills durable facts from raw dialogue, then reconciles them against the existing store via add / update / delete / noop. The most widely deployed reference for the extract-then-reconcile write path. Paper: Mem0.
Graphiti (Zep)30.2kApache-2.0Bi-temporal knowledge graph: edges carry validity intervals and are invalidated, not overwritten, so superseded facts remain queryable with time bounds. The closest production analogue to Memory Transactions, Versioning & Recovery.
cognee30.2kApache-2.0ECL (extract → cognify → load) pipeline turning documents and conversations into a typed knowledge graph plus vector index, with explicit prune/forget operations over the graph rather than append-only growth.
Letta24.4kApache-2.0The MemGPT lineage: OS-style paged memory where the agent edits its own core memory block through tool calls and pages the remainder to archival storage. Self-editing memory as the mechanism. Paper: MemGPT.
Hindsight20.9kMITConsolidation as an explicit policy layer over four levers — importance, merge, decay, eviction — instead of an unbounded append log. Industry framing already cited in Adjacent Research; paper: arXiv:2512.12818.
MemOS10.9kApache-2.0"Memory OS": a MemCube abstraction unifying plaintext, activation (KV), and parameter memory, with scheduling and cross-task reuse between them. Rare in treating KV-cache and long-term store as one substrate.
Honcho6.8kAGPL-3.0Memory as a representation of a person, not a transcript: background reasoning maintains per-peer models, queried through a dialectic API rather than similarity search.
MemMachine3.2kApache-2.0Clean two-tier service — episodic memory plus a durable user profile — with pluggable stores and a server/client split. A readable reference implementation if you are building your own.
MemoryOS1.6kApache-2.0EMNLP 2025 Oral ✅. Short / mid / long-term stores with heat-based promotion and eviction between tiers — one of the few systems whose admission and consolidation policy is inspectable code rather than a prompt.
redis/agent-memory-server308Apache-2.0Explicit working-memory (token-bounded, session-scoped) vs long-term memory split, with automatic extraction at the boundary. MCP + REST, Redis-native.

Paper → Code

Reference implementations for papers already indexed in this list — verified to exist and to correspond to the cited work.

Paper (in this list)CodeNotes
xMemoryHU-xiaobai/xMemory⭐120 · MIT. Author release for Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation.
AgeMemy1y5/AgeMem⭐37 · ACL 2026 SAC Highlight ✅. Unified long/short-term memory management learned as tool operations.
CraniMemPearlMody05/CraniMEM⭐7 · gated-and-bounded memory implementation; project page. ⚠️ No license file — all rights reserved by default; clear reuse with the authors.
HeLa-MemReinerBRO/HeLa-Mem⭐16 · ACL 2026 Main ✅. Hebbian associative distillation, reproduces the LongMemEval-S numbers. ⚠️ No license file — all rights reserved by default.
StructMemEvalyandex-research/StructMemEval⭐11 · Apache-2.0. Raw benchmark data + supplementary code from Yandex Research; authors label it work in progress.
memorywiremthamil107/memorywire⭐2 · Apache-2.0. Reference implementation of the wire format, including the provenance-based purge/quarantine path for poisoned stores.
A-MEMagiresearch/A-mem⭐1.2k · MIT. Zettelkasten-style memory that links and evolves notes on write. Paper reproduction lives in WujiangXu/A-mem (NeurIPS 2025 ✅). Last push Dec 2025 ⚠️.

Local-First & Coding-Agent Memory

Single-binary or file-backed systems, usually MCP servers, where the store is inspectable by a human and runs without a cloud dependency.

ProjectStarsLicenseWhat it actually implements
agentmemory27.3kApache-2.0Full lifecycle memory for coding agents: confidence scoring, decay and consolidation, knowledge graph, hybrid search, ~20 agent adapters. ⚠️ Headline ranking claims are self-reported.
Engram6.1kMITAgent-agnostic Go binary — SQLite + FTS5, MCP/HTTP/CLI/TUI, explicit conflict handling. Deterministic lexical recall as a deliberate counter-position to embedding everything (cf. ReFind in Retrieval & Recall).
memsearch (Zilliz)2.5kMITMarkdown files as the source of truth, Milvus for hybrid retrieval, typed episodic/procedural memory. Human-readable store = auditable memory.
Nocturne Memory1.3kMITRollbackable, visually inspectable graph memory server for MCP agents — edits are reversible operations, not blind writes. The practical counterpart to MemTX/ChronoMem-style recovery.
deja-vu1.0kMITNo write path: indexes the session transcripts 34 coding agents already write locally (lexical by default; optional deja embed may send text to an external embedding API) and serves them to any of the others — including pre-install history, so no cold start. Read-side mechanisms: outcome-aware ranking (a session that ended in a fix outranks one that gave up), redaction at ingest, per-project trust policy, forget/unforget. LoCoMo/LongMemEval drivers in-repo.
Vestige608AGPL-3.0Local-first Rust MCP server combining FSRS-6 decay, spreading activation, contradiction inspection, active suppression, Receipt Lock, and an inspectable dashboard. Built for causal recall — the quiet change that broke today, not the lookalike.
Mnemon507Apache-2.0Single Go binary, four-graph knowledge store with intent-aware recall, importance decay, and automatic deduplication. No API keys, one setup command.
Caura (ex-MemClaw)439Apache-2.0Governed shared memory for agent fleets: trust tiers, keystone policies, audit trails, multi-tenant scoping. The implementation counterpart to Memory Trust & Integrity and Multi-Agent Memory.
mengram189Apache-2.0Typed semantic / episodic / procedural memory where procedures are derived from recorded failures — one of the few OSS systems implementing the failure → procedure loop from Skill & Procedural Memory.
causal-memory66Apache-2.0Local-first Rust MCP server (also CLI + PyO3 bindings) where the store is one SQLite file holding facts, temporal state, and typed decision→outcome edges (caused / enabled / prevented). Spreading activation runs both ways — prevented edges carry negative activation, a GABA analogue — alongside Hebbian co-occurrence, Q-value dynamics, write-time gating, and immutable SWR-style consolidation. Reproducible benchmark harnesses (CausalEval, LoCoMo, LongMemEval) live in-repo; Hermes Agent provider plugin included.
Tree Ring Memory13MITProject-scoped Rust CLI lifecycle layer: SQLite/FTS explainable recall, deterministic consolidation, forgetting, redaction, audit, framework discovery. Memory that ages instead of accumulating raw transcript.
kgai1MITAppend-only, content-addressed decision log for dev teams: every change is a new event that supersedes the prior one through an explicit link with a reason, rejected paths stay queryable (temporal versioning), recall returns only decisions in force. Deterministic graph projection, lexical recall with no embeddings, and per-writer shard sync over the team's own S3 bucket (no server). Repo-supplied config is gated behind an explicit trust approval (governance). Claude Code plugin plus a Go CLI.
Hyperconsciousness1MITRust knowledge store with signed, encrypted append-only records and scoped, expiring grants for MCP access. Delegated grants cannot widen scope, add actions, or outlive their parent. Governance at the retrieval boundary; does not sandbox processes running as the owner. ⚠️ Developer alpha; no independent security audit.

Neuromorphic reference implementations — tinyHippo and Membrain — live in Neuromorphic & Bio-Inspired Memory.

Benchmarks & Evaluation Harnesses

ProjectStarsLicenseWhat it measures
LongMemEval1.0kMITICLR 2025 ✅. 500 questions across five long-term memory abilities over long interactive histories. The de-facto standard — and the number most vendor claims cite.
LongMemEval-V2132Apache-2.0Official successor, released 2026, addressing saturation and contamination in the original.
LoCoMo1.1ksee repo ⚠️Very long-term conversation benchmark (ACL 2024 ✅). Widely used dataset; frozen ⚠️ — last push Aug 2024, cite the data, don't expect maintenance.
MemoryData139none ⚠️Unified harness — 4 benchmark families, 22 method presets, one runtime — built specifically to make heterogeneous memory systems comparable. Directly attacks the fragmentation problem Evaluation & Benchmarks documents.
GateMem197MITMemory governance under multiple principals sharing one store: utility, access control, and active forgetting measured together instead of accuracy alone.
HaluMem155CC-BY-NC-ND-4.0 ⚠️Operation-level hallucination benchmark: scores extraction, updating, and answering separately, exposing errors that end-to-end QA accuracy hides. License is declared by README badge only — no LICENSE file in the repo; NC-ND terms forbid commercial use and derivatives, so treat it as read-only unless you clear it with the authors.
Mem-Gallery103MITACL 2026 Main ✅. Multimodal long-term conversational memory for MLLM agents. Last push Jan 2026 ⚠️.
STATE-Bench77MITMicrosoft. Memory-agnostic enterprise workflows measuring whether an agent improves with experience, not whether it recalls a fact. Blog write-up in Adjacent Research.
ResourceStarsWhy it's here
Agent Memory Techniques92630 runnable notebooks covering buffers, vector stores, KGs, episodic/semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, and LoCoMo evaluation. The best hands-on on-ramp before reading the papers above.
Awesome-AI-Memory (IAAR Shanghai)1.2kBroad bilingual knowledge base spanning LLM memory and agent memory, including engineering frameworks and applications.
Awesome-Agent-Memory (TeleAI)596Systems + benchmarks + papers for LLM and MLLM memory — stronger multimodal coverage than this list.
Awesome-GraphMemory339Survey-grade depth on graph-based agent memory specifically.
Awesome-Agent-Memory-Papers229Paper-only index with a browsable website.
OpenDataBox/awesome-agent-memory54Same name, different list — a four-axis paper taxonomy, companion to the MemoryData harness above.

Adjacent Research

These papers and projects aren't solely about agent memory but contribute relevant architectural ideas:

Paper / ProjectDateRelevance
MemTools: A Unified Research Framework for Interoperable Agent Memory — Zhao, Chen, Liang, He, Wang, Zhao, LiuJul 2026Targets architectural fragmentation as the thing blocking systematic memory research: implementations couple lifecycle stages together, entangle evaluation logic with specific datasets, and barely support heterogeneous memory types. MemTools standardizes the memory lifecycle through declarative data contracts, enabling interchangeable assembly of components across systems, and orthogonally separates benchmark datasets from execution protocols so evaluation can be reconfigured independently of data. Adds a unified computational interface for coordinating symbolic, neural, and multimodal representations in one runtime. ⚠️ Authors label it work in progress; the evaluation demonstrates isolation of design variables rather than task-level performance.
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories — Lovén, Sauvola, Riekki, TarkomaAug 2026Agents share transport and can call each other's tools, but have no protocol for reconciling what they know. MELD admits every incoming claim through a five-outcome procedure (insert, merge, relate, conflict, reject) decided from three signals — scoped claim-key identity, embedding similarity, and a natural-language- inference verdict — under context and freshness gates, acting through exactly one auditable, authenticated Patch as the sole object that mutates state, with a per-claim status CRDT over standard pub/sub. The design stance is the notable part: MELD does not adjudicate truth — a detected contradiction is preserved for later adjudication, never silently resolved. Merge classifier separates at AUC 0.968 with a 0.013 false-merge rate; the status CRDT reconverges in 30/30 real partition-heal trials where last-writer-wins manages 11/30; recall-non-inferior to a centralized store at ~11% less live storage; ~3× fewer messages at matched recall. Evaluated on a real continuum spanning operator-grade 5G edge, national HPC, and a local tier 📄. ⚠️ Preprint; code and data on Zenodo.
memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations — MunirathinamMay 2026 (rev Jul 2026)JSON-Schema 2020-12 wire format for 5 memory ops (remember/recall/forget/merge/expire) × 4 memory types over mem0/Letta/Cognee/Zep/pgvector backends, with an optional HITL governance channel. Not a new algorithm — a packaging of RRF/FSM/consolidation into a vendor-neutral protocol meant to compose with MCP. v3 adds a poison-recovery eval (PurgeBench) showing its provenance field is the strongest lever for recovering a poisoned store — relevant complement to TMA-NM/GhostWriter above.
EverMemOS — Liu, Bai, Chen et al.Jan 2026Self-Organizing Memory Operating System for structured long-horizon reasoning.
AriGraph — Anokhin, Radionov et al.IJCAI 2025 ✅Knowledge Graph World Models with episodic + semantic memory. Proceedings
Memory Matters More Than You Think — Zhou, Li et al. (Renmin U.)Jan 2026Event-centric memory with Logic Map for agent searching and reasoning.
The Consolidation Problem in Agent Memory — Vectorize (Hindsight)May 2026Industry framing of consolidation as a policy layer with four levers — importance, merge, decay, eviction — benchmarked across Mem0, Zep, Letta, LangChain, Hindsight.
Introducing STATE-Bench — Microsoft (repo)May 2026Open-source, memory-agnostic benchmark measuring whether agents improve with experience on stateful enterprise tasks (travel, support, shopping) — not just recall.
State of AI Agent Memory 2026 — Mem02026Industry survey of benchmarks, architectures, and six open problems (temporal abstraction, cross-session structure/identity, app-level eval, privacy/consent, staleness).

Key Themes & Cross-Cutting Insights

These papers, taken together, converge on several architectural principles:

1. 🚪 Gated Admission — Not Everything Deserves to be Remembered

Papers: A-MAC, CraniMem, SleepGate, ACC

The strongest systems filter before storing. CraniMem uses goal-conditioned cosine gating; A-MAC scores across 5 dimensions (utility, confidence, novelty, recency, type). The alternative — storing everything and retrieving selectively — leads to noise accumulation and proactive interference.

2. 📦 Typed Multi-Store — One Vector Store Is Not Memory

Papers: xMemory, SYNAPSE, AdaMem, CraniMem, Mem0, Episodic Memory

Every serious architecture separates memory by type (episodic, semantic, procedural, working). xMemory's 4-level hierarchy shows 44.91% of memories get retroactively reassigned — proving flat structures fundamentally fail. The human brain doesn't use one store; neither should your agent.

3. 🔄 Consolidation Loops — Memory Must Evolve

Papers: SleepGate, CraniMem, SYNAPSE, A-MEM

Raw episodic traces should consolidate into durable semantic knowledge over time. SleepGate's sleep cycles, CraniMem's replay-based promotion, and SYNAPSE's spreading activation all implement variants of this. Without consolidation, memory becomes a write-only log.

4. 🗑️ Forgetting as a Feature — Not a Bug

Papers: SleepGate, Procedural Memory, xMemory

SleepGate reduces proactive interference from O(n) to O(log n) through learned forgetting. This is the most under-implemented capability in agent frameworks — most systems only add memories, never remove or decay them. Forgetting is what keeps memory useful rather than just large.

5. 📊 RAG ≠ Agent Memory

Papers: xMemory, Memory Survey, Anatomy of Agentic Memory

The clearest thesis across this literature: retrieval-augmented generation is a retrieval strategy, not a memory system. Agent memory requires admission, evolution, consolidation, forgetting, typed storage, and context-aware retrieval. RAG addresses only the last item.

6. 🧪 Evaluation Is Broken

Papers: Anatomy of Agentic Memory, StructMemEval

Current benchmarks are saturated, metrics misalign with semantic utility, and results vary wildly by backbone model. StructMemEval reveals agents can't organize memory autonomously. The field needs better evaluation before it can claim progress.


7. 🎮 RL Is Replacing Heuristic Memory Management

Papers: AgeMem, EMPO², MemFactory, Memory-R1

The dominant shift in 2026: from hand-coded admission/retrieval rules to reinforcement-learned memory policies. AgeMem exposes all memory operations as tool-based actions and learns when to use them via GRPO. EMPO² takes this further by using memory for exploration, not just recall. MemFactory provides the infrastructure to train these policies. This is the most consequential paradigm shift since the field recognized that RAG ≠ memory.

8. 🔒 Memory Privacy Becomes a First-Class Concern

Papers: ADAM, FSFM

April 2026 surfaced the first systematic memory-extraction attack (ADAM, up to 100% ASR) — and the first forgetting frameworks (FSFM) that explicitly frame selective forgetting as a privacy primitive, not just a retention/efficiency one. Expect memory threat-modeling, audit interfaces, and unlearning APIs to follow.

9. 🪜 Memory ↔ Skills ↔ Rules as a Compression Spectrum

Papers: Experience Compression Spectrum, Externalization in LLM Agents

The memory and skill-discovery communities have been solving the same problem (extract reusable knowledge from interaction traces) without citing each other — cross-community citation rate is below 1%. The Compression Spectrum reframes both as points on a single axis (5–20× → 50–500× → 1000×+ compression), and identifies the missing diagonal: no current system supports adaptive cross-level compression.

10. 🧱 Active, Constructive Memory Operations

Papers: MemReader (active extraction), SCG-MEM (schema-constrained generation), HeLa-Mem (Hebbian distillation)

Two paradigm shifts in April 2026: (1) MemReader moves extraction from passive transcription to an RL-learned WRITE/DEFER/RETRIEVE-CONTEXT/DISCARD policy. (2) SCG-MEM replaces dense retrieval with schema-constrained decoding (Piaget-style assimilation vs accommodation). Together they suggest the future of memory is generative-constructive, not retrieve-and-paste.

11. 🧩 Reconstruction ≠ Retrieval — Memory Access as Active Reasoning

Papers: MRAgent, MemPro, DeltaMem

Wave 6's headline shift: stop treating memory access as a static retrieve-then-reason step. MRAgent folds LLM reasoning into graph traversal (Cue-Tag-Content + active reconstruction), iteratively pruning paths on accumulated evidence. MemPro generalizes the idea to the whole pipeline — the memory system itself becomes an evolvable program that rewrites its own construction/retrieval logic. DeltaMem reconstructs each experience on demand by composing a root-to-leaf residual chain rather than storing flat copies. Memory is increasingly constructed at access time, not fetched.

12. 📉 Memory Economics Gets Measured

Papers: Agent Memory Characterization, Beyond the Context Window, Personalize-then-Store

For two years memory papers reported accuracy; few reported cost. Wave 6 changes that. The first systems characterization (Omri et al.) builds a phase-aware cost profiler and a 4-axis taxonomy, deriving 10 deployment recommendations across the write and read paths. Personalize-then-Store reframes admission as a budget problem — gating transient sessions to spend the memory budget where it pays off. The question is shifting from "does memory help?" to "what does memory cost, and where does it amortize?"

13. 🔐 Trust & Integrity Become Explicit, Distinct from Privacy

Papers: TRUSTMEM, MemSyco-Bench, A-TMA

Wave 6 established memory as an attack surface (ADAM). Wave 7 shows a second, orthogonal failure class: memory that's technically retrieved correctly but still misleads. TRUSTMEM targets self-inflicted corruption during consolidation (omission, hallucination). MemSyco-Bench targets over-reliance — agents favoring a user's stored prior view over current evidence. A-TMA targets ghost memory — stale and current facts silently coexisting in the same bank. None of these are extraction attacks; they're reliability failures baked into ordinary write/retrieve operations.

14. 🔬 Evaluation Moves from "Did the Task Succeed" to "What Does the Memory Actually Contain"

Papers: MemProbe, Are We Ready For An Agent-Native Memory System?

MemProbe's headline finding — task completion nearly saturates even for a memoryless baseline, while structured state recovery from the memory artifact itself stays around 0.6 — is a direct rebuke of end-to-end accuracy as a memory metric. "Are We Ready" pushes the same point systemically: decomposing 12 memory systems into 4 modules shows no architecture dominates, and effectiveness is bottleneck-specific, not a single leaderboard number. Both extend Anatomy of Agentic Memory's Feb-2026 critique from theory into large-scale measurement.

15. 🧮 Compaction Gets a Unifying Theory

Papers: What to Keep What to Forget (Rate–Distortion View), Is Agent Memory a Database?

Two Wave 7 papers step back from architecture-of-the-week and ask for formal foundations. The rate–distortion paper shows KV-cache eviction, prompt pruning, bounded recurrent state, and agent memory consolidation are the same budget-constrained retain/discard decision, and that all four fail identically (irreversible, pre-query, recency/attention-biased). "Is Agent Memory a Database?" argues memory correctness is a property of the state trajectory, not individual records — no record-level storage system can satisfy it, motivating the GEM formalization. Both suggest the field's next gains come from theory, not another heuristic gate.

16. 🔑 Authority Is a Separate Property from Content — and Consolidation Destroys It

Papers: AuthMem-Bench, PPMF

Two independent papers landed the same finding within days, from opposite directions. Summarization is lossy in a way nobody was measuring: it preserves the claim and drops the permission. AuthMem-Bench measures it as an accident — authority collapse in 48 of 49 configurations, a 50.3% mean unauthorized-action rate for collapsed memories. PPMF weaponizes it — laundering an untrusted observation into apparent user history, preserving the trigger while erasing the low-trust source, up to 1.000 ASR. Both fixes are the same shape: persist authority as first-class metadata and match action risk against it at tool-call time. This is distinct from insight #13, which asks whether a memory is true; this asks what a memory is allowed to authorize.

17. 🧾 A Memory Write Is Not a Belief Commit

Papers: MemTX, MemTxn, ChronoMem, TARL

The transactions section in one line. Agent memory is acquiring database semantics — staged writes, snapshot isolation, versioned heads, rollback, and cascading repair of derived records and tool side effects. The forcing function is that memory now drives irreversible external actions, so the cost of an unsound write is no longer a bad answer but a real-world consequence. Note the direction of travel: TARL replaces a binary Write/Hold with five actions; ChronoMem adds an undo; MemTX gates tool calls on in-flight belief state. All three treat the moment of writing as provisional by default.

18. 🧱 Structure Is Not Free — and May Not Be Necessary

Papers: ReFind, Filesystem-Based Memory, Harness the Memory, Keep It InMind

The sharpest contrarian thread this list has carried. ReFind beats HippoRAG 2 (58.2 vs 53.2) with no semantic structure at all, on a matched backbone. Filesystem-Based Memory finds organization buys retrieval economy, not accuracy — and erodes as stores grow. Harness the Memory finds no substrate dominates and that excessive retrieval actively harms sequential decision-making. Against all three, Keep It InMind argues the framing itself is wrong: the failure isn't structure versus no-structure but the query-conditioned retrieval interface — 84.0% with memory in context collapsing to ≤14.4% when the same memory must be retrieved. Read together they pose an uncomfortable question for most of this list: is the field over-investing in write-time structure to compensate for a broken read-time interface?

Open Questions

  • How should admission gates interact with consolidation loops? (A-MAC + SleepGate integration)
  • What's the right forgetting curve for different memory types? (No paper addresses type-specific decay)
  • Can skill routing be unified with memory retrieval? (SkillRouter + xMemory convergence)
  • What's the minimum viable memory architecture? (Most papers add complexity — none simplify)
  • How should backward verification loops (MemMA probe-verify-repair) integrate with forward memory policy? (MemMA + AgeMem convergence)
  • Can RL-learned memory policies transfer across domains? (EMPO² shows OOD promise — but limited evaluation)
  • Can constructivist memory (schema-bound decoding, SCG-MEM) replace similarity-based retrieval?
  • Can the "missing diagonal" be filled — adaptive cross-level compression between episodic/procedural/declarative?
  • How robust is agent memory against systematic extraction attacks (ADAM-class threats)?
  • What is the right interface between active extraction (MemReader) and admission gating (A-MAC)?
  • Does Hebbian distillation (HeLa-Mem) outperform LLM-driven semantic abstraction at scale?
  • Can access-time reconstruction (MRAgent) scale without latency blowup versus precomputed retrieval?
  • What is the right cost model for agent memory at fleet scale, and when does a memory system amortize its write cost? (Agent Memory Characterization)
  • Do trust verifiers (TRUSTMEM-style) generalize across consolidation architectures, or must every memory system train its own?
  • Is memory-induced sycophancy (MemSyco-Bench) an architectural property or a training-data artifact inherited from the base LLM's own sycophancy?
  • Does ghost memory (A-TMA) require an architectural fix, or is decoupled bank/retrieval/answer-level evaluation enough to catch it before deployment?
  • Can a single rate–distortion budget metric make KV-cache eviction, prompt compression, and agent memory consolidation apples-to-apples comparable?
  • If task success saturates even for memoryless baselines (MemProbe), how many published memory-system benchmark wins are actually measuring memory at all?
  • If an agent with controllable lexical search over raw logs (ReFind) matches graph memory, which workloads actually justify the cost of write-time structure?
  • Can authority metadata (AuthMem-Bench, PPMF) survive consolidation without a full transactional substrate (MemTX), or are the two the same requirement discovered from different ends?
  • What is the right granularity for memory rollback — whole-store snapshots (ChronoMem) or per-record validity intervals (MemTX, TARL)?
  • Does native in-backbone memory (Metis) subsume external memory systems, or does it just relocate the same admission, forgetting, and provenance problems inside the weights where they can no longer be audited?
  • If retrieval fails on implicit associations (Keep It InMind) even at 8× embedding dimensionality, is routing-by-visibility a fix or an admission that the retrieve-on-query interface is the wrong abstraction?

Contributing

This list grows as the field grows. To contribute:

  1. Open an issue or PR with the paper title, arXiv/DOI link, authors, date, and a 1-2 sentence summary of the key contribution
  2. Suggest which section it belongs in (or propose a new one)
  3. Bonus: note how it relates to or challenges existing papers in the list

Papers should be peer-reviewed, at reputable workshops (e.g., ICLR MemAgents), or highly cited preprints. We prioritize papers with novel architectural ideas over incremental benchmark improvements.

For repositories (the Implementations section), the bar is different: a repo must implement a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance — and either back a paper in this list, show real adoption, or fill a niche nothing else covers. Star counts are context, not a criterion. Self-reported benchmark rankings are not evidence; a reproducible harness is. Active maintenance and an OSI-approved license are required only in the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where the code is meant to be vendored. Author code for papers and benchmark/dataset repos is judged on reproduction value instead: a license must still be disclosed, and where it is absent or non-OSI (e.g. NC-ND, or no LICENSE file at all) — or where the repo has been quiet for more than six months — the row is flagged ⚠️ rather than dropped. A frozen benchmark that everyone cites still earns its place; it just gets its last-push date stated.


License

CC0

This list is released under CC0. The papers themselves retain their original licenses.

agentic-ai
agent-memory
awesome-list
cognitive-architecture
llm-agents
memory-systems
neuromorphic

tfatykhov/awesome-agent-memory

Curated research on memory systems for LLM agents

23

25 commits

updated Sep 27, 2026

See the code

README

🧠 Awesome Agent Memory

Awesome

A curated collection of research papers on memory systems for LLM-based agents — covering architecture, retrieval, forgetting, consolidation, evaluation, the cognitive science that inspires it all, and the neuromorphic hardware that may one day implement it.

Why this list? Every agent framework bolts on a vector store and calls it "memory." These papers show what memory actually requires: admission control, consolidation loops, forgetting mechanisms, typed multi-store architectures, and retrieval strategies that go far beyond cosine similarity.

Maintained by @tfatykhov · Built alongside Nous, a cognitive AI agent with typed multi-store memory.

🚀 Start Here

If you want to...Start with
Understand the landscapeMemory Survey (taxonomy) → Anatomy (critical analysis)
Build agent memoryxMemory (why RAG isn't enough) → A-MAC (admission control) → Mem0 (production system)
Add forgetting/consolidationSleepGate → CraniMem
Evaluate your memory systemStructMemEval → Anatomy (why benchmarks are broken)
Train memory with RLAgeMem (tool-based ops) → EMPO² (exploration) → MemFactory (framework)
Understand memory privacy riskADAM (extraction attack) → FSFM (forgetting as defense)
Unify memory ↔ skills ↔ rulesExperience Compression Spectrum → Externalization
Make memory writes safe & reversibleMemTX (transactional belief commit) → ChronoMem (versioning + rollback)
Ship something this weekImplementations & Reference Code (curated repos, mechanism-by-mechanism)
Explore the neuroscienceNeuromorphic section

Contents

Verification Legend

SymbolMeaning
✅Venue confirmed — verified via official proceedings, OpenReview, or arXiv metadata
📄Self-reported — metrics from authors' own evaluation; exercise caution
⚠️Unverified — plausible claim but not independently confirmed
🔬Editor's synthesis — cross-paper connection identified by the maintainer, not claimed by original authors

Surveys & Taxonomies

PaperAuthorsDateKey Contribution
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory ProblemTangAug 2026The first serious attempt to say what a memory is, formally: memory is a basis, knowledge is its span, and answerability is a coverage problem — a query is answerable exactly when some single item in the span covers it. Optimal memory becomes the capacity-constrained maximizer of expected coverage, tracing a utility–capacity frontier that serves as a common yardstick for comparing memory systems. Also treats noise (coverage vs precision when the store holds false claims) and formalizes continual memory as a sequential MDP — memory is the state, writing is the action, query-time utility is the delayed reward. Instantiated concretely on Homer's Odyssey. ⚠️ Preprint; solo-author.
Memory in the Age of AI Agents: A Survey47 authorsDec 2025Unified taxonomy across 3 lenses: Forms (token/parametric/latent), Functions (factual/experiential/working), Dynamics (formation/evolution/retrieval). Distinguishes agent memory from RAG and context engineering. Hugging Face Daily Paper #1 ⚠️ (unverifiable — HF archive page returns 404).
Anatomy of Agentic MemoryJiang, Li, Wei et al.Feb 2026Structured taxonomy of Memory-Augmented Generation (MAG) systems. Critically analyzes empirical fragility: benchmark saturation, metric misalignment, backbone-dependent variance, overlooked latency costs. Evaluates 5 systems (LOCOMO, A-Mem, MemoryOS, Nemori, MAGMA).
The Landscape of Agentic RL for LLMsGuibin Zhang et al. (Oxford, Shanghai AI Lab, NUS, UIUC, UCL)Sep 2025Synthesizes 500+ works. Reframes LLMs as autonomous agents using POMDPs. Covers planning, tool use, memory, reasoning, self-improvement.
From Static Templates to Dynamic Runtime GraphsYue et al. (IBM Research)Mar 2026Agentic Computation Graphs (ACGs) framework — distinguishes workflow templates, realized graphs, and execution traces. Organizes ~40 papers by when structure is determined.
Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging FrontiersPengfei DuMar 2026Write–manage–read loop formalization. 3D taxonomy: temporal scope × representational substrate × control policy. Five mechanism families: context-resident compression, retrieval-augmented stores, reflective self-improvement, hierarchical virtual context, policy-learned management. Proposes vision of a "foundation model for memory control."
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness EngineeringZhou, Chai, Chen, Guo, Shan, Song, Xu et al.Apr 2026Unified review across four externalized components: memory stores, reusable skills, interaction protocols, runtime harness. Argues modern agent capability comes from reorganizing the runtime around the model, not from changing weights. Connects memory literature with skill-discovery and protocol-design literatures that rarely cite each other.
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM AgentsZhang, Wang, Cui, Qiu, Li, Zhu, HeApr 2026Positions memory/skills/rules on a single compression axis (5–20× episodic, 50–500× procedural, 1000×+ declarative). Citation analysis of 1,136 references finds <1% cross-community citation between memory and skills literatures. Identifies the "missing diagonal" — no current system supports adaptive cross-level compression.
From Storage to Experience: A Survey on the Evolution of LLM Agent Memory MechanismsLuo, Tian, Cao, Luo et al.May 2026ACL 2026 Findings ✅ — Maps the field's evolution from passive storage stores toward experience-centric, self-evolving memory. Companion paper-list: Evolving-LLM-Agent-Memory-Survey.
Graph-based Agent Memory: Taxonomy, Techniques, and ApplicationsYang, Zhou, Xiao, Dong et al.Feb 2026First focused taxonomy of graph-based agent memory: node/edge designs, construction strategies (entity-centric vs event-centric vs hybrid), retrieval (one-hop vs multi-hop vs subgraph), and update/forget operators. Useful companion to A-MEM, HeLa-Mem, GAM.
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent MemoryOrogat, MansourMay 2026Argues record-level correctness (rows, embeddings, edges) cannot satisfy long-term memory's needs, causing four recurring failure modes: unregulated growth, missing semantic revision, capacity-driven forgetting, read-only retrieval. Formalizes Governed Evolving Memory (GEM) — state-level operators (ingestion, revision, forgetting, retrieval) with six correctness conditions on the trajectory, not individual records. Prototype MemState on a property-graph backend. Proves no record-level system can satisfy the conditions regardless of storage model. ⚠️ Preprint.

Memory Architectures

PaperAuthorsDateKey Contribution
Metis: Memory Foundation ModelZhang, Guo, Sun, Zhang, Hao, Lin et al. (17 authors)Jul 2026Challenges this list's prevailing external-module framing. Argues agent memory should be native to the backbone, and formalizes native memory as (a) a persistent, dynamically evolving memory state inside the model and (b) native memory procedures that store and use information through model computation. Metis equips a foundation model with a native memory state accessed via memory attention, acquired through mid-training on large-scale memory-specific data. Online memory maintenance is gradient-free — a memory update requires only a forward pass, and all learned weights remain frozen at inference. Project and model checkpoints released. ⚠️ Preprint.
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and SustainabilityZhou, Yu, Wei, Wu, Ouyang, Jiao et al. (11 authors)Jul 2026The first systematic study of the memory design that is actually deployed by default — a directory tree of markdown files the agent itself reads, writes, and reorganizes with generic file tools — which prior research largely passed over. Formalizes it as three roles around one memory filesystem (management, search, execution agents), unifying declarative memory and skills in a single store. The default's two working assumptions do not hold up: what organization reliably buys is search economy (organized stores roughly halve retrieval cost where material is large), not better answers — no agent measured converts organization into accuracy, and organization erodes as the store grows for all but the strongest management agent. Changing the tool set alone reshapes the store as strongly as swapping the model 📄. ⚠️ Preprint.
MAGE: Memory as Agent-Guided ExplorationChen, Lai, Feng, Han, Zhang, Lu, Li et al.Jun 2026Argues semantic-similarity organization mismatches execution-state dependencies on long-horizon tasks — it fragments decision trajectories and mixes valid/erroneous traces. Stores interactions in a hierarchical state tree; agent state derives from the active root-to-current path (subgoal summaries + recent traces + branch hints). Four coupled ops: Grow, Compress, Maintain, Revise. +7.8–20.4pp task success on MemoryArena, -55.1% tokens 📄.
Mechanistic Attention Guidance for Agent Memory Refinement (AGMR)Hong, Qiu, Wang, YaoJul 2026Existing self-evolving memory only inspects textual outputs (trajectories, reflections), risking unreliable error attribution and hallucinated edits. Uses retrieval-head attention as a mechanistic signal — aggregates attention over memory segments × decision steps into a context-utilization matrix that reveals actual memory-use patterns. AGMR corrects/enhances memory on failure, simplifies it on success, and re-executes to verify each update. Beats text-only refinement baselines on interactive decision-making benchmarks. Code released. ⚠️ Preprint.
Supra Cognitive Modes (SCM): A Routed Architecture for Agent MemoryTobkin, YangJul 2026Routes each query — factual lookup, relation-chain/current-state reasoning, or broad long-history synthesis — to a matched retrieval+synthesis payload over one shared ingest substrate (multi-granularity embeddings, extracted triples, fact-version metadata). A frozen semantic classifier + runtime gates dispatch among fused lexical/dense lookup, graph/multi-hop, and stratified long-form synthesis. Reports 84.87% LoCoMo factoid / 68.61% adversarial abstention, 61.49% MAB, 86.00% LongMemEval — but the authors themselves flag causal routing effects and efficiency gains as outside the available evidence 📄 (unusually candid self-limitation disclosure). ⚠️ Preprint.
A-MEM: Agentic Memory for LLM AgentsXu, Liang, Mei, Gao, Tan, ZhangFeb 2025Dynamic memory organization inspired by Zettelkasten. Creates interconnected "memory notes" with content, metadata, and explicit links. Agent autonomously decides when to create, update, or link memories. NeurIPS 2025 ✅.
SYNAPSE: Episodic-Semantic Memory via Spreading ActivationJiang, Chen, Pan et al.Jan 2026Memory as a dynamic graph where relevance emerges from spreading activation rather than pre-computed vector links. Lateral inhibition and temporal decay highlight relevant sub-graphs. Triple Hybrid retrieval.
MAGMA: Multi-Graph Agentic Memory ArchitectureJiang, Li, Li, LiJan 2026Multi-graph based architecture separating different memory concerns into distinct graph structures for AI agents. ACL 2026 Main ✅.
Mem0: Production-Ready AI Agents with Scalable Long-Term MemoryChhikara, Khant, Aryan, Singh, YadavApr 2025Scalable memory-centric architecture with dynamic extraction, consolidation, and retrieval. Graph-based variant for complex relational structures. 26% improvement over OpenAI memory, 91% lower latency, 90%+ token savings 📄 (self-reported). Evaluated on LOCOMO.
CraniMem: Neurocognitively Motivated Gated & Bounded MemoryMody, Panchal, Kar, Bhowmick, KaraniMar 2026Cranial-inspired dual-store: bounded FIFO episodic buffer + knowledge graph (long-term). Goal-conditioned gating. Utility tagging (Importance + Surprise + Emotion). Beats Mem0 by +57.6% on noisy HotpotQA 📄 (self-reported; "noisy" = authors' own distractor injection, not standard benchmark). ICLR 2026 MemAgents Workshop ✅. Code
AdaMem: Adaptive User-Centric Memory—Mar 2026Working, episodic, persona, and graph memories. Key innovation: question-conditioned retrieval planning — resolves target participant first, builds retrieval route combining semantic + relation-aware graph expansion. SOTA on LoCoMo and PERSONAMEM.
PlugMem: A Task-Agnostic Plugin Memory ModuleYang, Galley, Wang, Gao, Han, Zhai (Microsoft Research, UIUC)Mar 2026Converts raw interactions into propositional (facts) + prescriptive (skills) knowledge, organized in a knowledge-centric memory graph. Knowledge — not entities or chunks — is the unit of memory access. Single plug-and-play module outperforms task-specific designs across 3 diverse benchmarks. Highest information density: more useful info per context token consumed 📄.
Memora: Harmonic Memory RepresentationXia, Zhang, Dixit, Harimurugan, Wang, Ruhle, Sim, Bansal, Rajmohan (Microsoft Research)Feb 2026Proves that RAG and KG memory are special cases of this unified framework. Primary abstractions index concrete memory values; "cue anchors" expand retrieval beyond semantic similarity. New SOTA on both LoCoMo AND LongMemEval 📄. ICML 2026 ✅.
MemMA: Coordinating the Memory Cycle through Multi-Agent ReasoningLin, Zhang, Lu, Liu, Tang, He, Zhang, Wang (Microsoft Research)Mar 2026Multi-agent framework: Meta-Thinker → Memory Manager → Query Reasoner. Backward path innovation: synthesizes probe QA pairs, verifies memory, converts failures into repairs BEFORE finalizing. Plug-and-play — improves 3 different storage backends on LoCoMo. Code
H-Mem: Hybrid Multi-Dimensional MemoryYe, Huang, Chen, Zhang (Rutgers)Mar 2026Organizes memory across time AND topic dimensions simultaneously. Mimics associative + hierarchical properties of human memory. EACL 2026 ✅.
H-MEM: Hierarchical Memory for High-Efficiency Long-Term ReasoningSun et al.Mar 2026Multi-level memory storage with positional index encoding of sub-memory at each layer. Confidence-weighted retrieval — attaches memory weights to provide LLMs with uncertainty reference. Key finding: retrieval is ineffective without structured hierarchical storage. EACL 2026 ✅.
GAM: Hierarchical Graph-based Agentic Memory for LLM AgentsWu, Zhang, Lin, Xu, Xu, Chen, Zou et al.Apr 2026Explicitly decouples encoding from consolidation: an event-progression graph captures stream updates online; integration into the topic-associative network is deferred until a semantic shift is detected. Addresses the tension between stream-based fluidity and structured retention.
HeLa-Mem: Hebbian Learning and Associative Memory for LLM AgentsZhu, Li, Zhang, Liu, YangApr 2026ACL 2026 ✅ — Bio-inspired dual-graph: (1) episodic graph evolves via Hebbian co-activation; (2) semantic store populated via Hebbian Distillation — a Reflective Agent identifies densely-connected hubs and distills them into reusable semantic knowledge. Beats prior SOTA on LoCoMo across 4 categories with fewer context tokens 📄. Code
LightMem: Lightweight LLM Agent Memory with Small Language ModelsZhang, Zhang, Chen, Huang, Zheng et al.Apr 2026ACL 2026 ✅ — SLM-driven memory with strict online/offline separation. STM/MTM/LTM tiers with two-stage retrieval (vector coarse → semantic re-rank). ~2.5 F1 over A-MEM on LoCoMo, 83ms retrieval, 581ms end-to-end 📄. Shows that careful SLM use can replace repeated large-model memory calls.
MemMachine: Ground-Truth-Preserving Memory for Personalized AI AgentsWang, Yu, Love, Zhang, Wong, Scargall, Fan et al.Apr 2026Open-source system integrating short-term, long-term episodic, and profile memory in a ground-truth-preserving pipeline. Targets multi-session degradation in standard RAG.
Omni-SimpleMem: Autoresearch-Guided Lifelong Multimodal MemoryLiu, Ling, Qiu, Liu, Han, Xia, Tu et al.Apr 2026Uses autonomous research-agent search over the design space (architecture × retrieval × prompts × data pipeline) to discover effective lifelong multimodal memory configurations. The first paper to treat memory architecture itself as a search target.
Human-Inspired Memory Architecture for LLM AgentsKerestecioglu, Robsky, Vasters, Sharma, Kesselman (Microsoft)May 2026Biologically-grounded architecture with six cognitive mechanisms: sleep-phase consolidation, interference-based forgetting, engram maturation, reconsolidation on retrieval, entity knowledge graphs, hybrid multi-cue retrieval. Introduces a synthetic calibration methodology that derives all thresholds without benchmark exposure — eliminates a common eval-leakage source. First streaming M-tier LongMemEval eval (475 sessions). Dedup-based consolidation: 97.2% retention precision, 58% store reduction (+21.8 pp) on a 13K-issue VSCode dataset 📄.
MRAgent: Memory is Reconstructed, Not Retrieved — Graph Memory for LLM AgentsJi, Li, Hooi (NUS)Jun 2026ICML 2026 ✅ — Represents memory as a Cue-Tag-Content graph where associative tags bridge fine-grained cues to contents. An active reconstruction mechanism folds LLM reasoning into memory access — iteratively exploring and pruning retrieval paths on accumulated evidence, adapting retrieval to the reasoning context while avoiding combinatorial blowup. Up to 23% over strong baselines on LoCoMo + LongMemEval with substantially lower token/runtime cost 📄. The cleanest articulation of the "reconstruction ≠ retrieval" paradigm.
MemPro: Agentic Memory Systems as Evolvable ProgramsLiu, Wang, Wu, Huang, Tao, Song, Zhou, He (ECNU)May 2026Treats the entire memory construction-retrieval (MCR) pipeline as an evolvable program, not just the memory bank. Maintains a version tree of runnable memory-system implementations; an Evolving Agent selects promising versions, diagnoses recurring failures, and synthesizes improved children via failure-mode-guided edit-debug refinement. Beats static and prompt-level evolving baselines on LongMemEval / LoCoMo / HotpotQA / NarrativeQA within a few iterations, with favorable performance-cost trade-off 📄. Code available. ⚠️ Preprint; venue unconfirmed.
MemDreamer: Hierarchical Graph Memory with Agentic Retrieval for Long Video—Jun 2026Extends agent memory to the multimodal long-video setting: a hierarchical graph memory plus agentic retrieval over hours-long video, exploiting the tree+graph structure for question answering. ⚠️ Preprint; venue unconfirmed.
H-Mem (Yu & Fang): Evolving & Retrieving Agent Memory via a Hybrid StructureYu, Fang, Liu, Ma (CUHK-Shenzhen / Huawei)May 2026Distinct from the EACL H-Mem (Rutgers) and the Sun et al. H-MEM listed above — a third system sharing the name. Hybrid tree + graph: a temporal-semantic tree lets short-term memory evolve into a summarizing long-term store, while a parallel knowledge graph captures entity relationships; retrieval exploits both. SOTA on QA across three agent-memory benchmarks 📄. ⚠️ Preprint; venue unconfirmed.
MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context ManagementLiu, Wu, Liu, Zhao, Liu, Li, Zhang, Wang, Guo, LiuJun 2026Extends agent memory to long-horizon mobile GUI agents. Introduces Context-as-Action (ConAct): context management becomes a first-class action emitted by the same policy that selects UI actions, maintaining three structured fields (folded action history, folded UI state, recent step record) instead of passively appending ReAct-style transcripts. 8B model trained on the 2,956-trajectory MemGUI-3K dataset achieves best open-data 8B performance on MemGUI-Bench and generalizes OOD to MobileWorld. Code/data to be released. ⚠️ Preprint.
Multi-Head Recurrent Memory AgentsLi, Yeh, LiJul 2026Diagnoses why recurrent memory agents degrade at long context: decomposing performance into capture vs retention shows retention collapse is the dominant bottleneck, caused by treating memory as one monolithic text block that every update risks overwriting. Multi-Head Recurrent Memory (MHM) partitions memory into independent heads with a stage-wise select-then-update strategy — only one head updates per step, the rest are structurally shielded. The MHM-LRU instantiation is training-free and lifts RULER-HQA retention at 896K tokens from <30% to 73.96%. Architecture, not model behavior, is the lever. 📄 ⚠️ Preprint.

Memory Admission & Gating

PaperAuthorsDateKey Contribution
When Not to Write Memory: Governing False Promotion from Correlated Agent Traces (GovMem)Qi, Xu, LiJun 2026Reframes admission as a write-path governance problem: repeated observations aren't independent evidence if copied from a shared source, induced by a shared prompt, or valid only in a narrower scope. GovMem estimates dependency-aware support, retrieves counterevidence, assigns scope, and outputs promote / reject / needs-review. Cuts false promotion 0.597→0.040 (synthetic) at 0.960 recall; on a human-labeled real-trace subset, false promotion falls to 0.032 but held-out false promotion stays 0.111 at 0.692 review burden. On 133 high-impact external coding-agent candidates, none were judged safe for automatic promotion. Positioned as a diagnostic governance design point, not a validated auto-writer — a sobering result for anyone auto-promoting agent traces into memory 📄. Accepted at MLISE 2026 ✅.
A-MAC: Adaptive Memory Admission ControlZhang et al. (Workday AI)Sep 20255-dimension scoring: Utility (LLM call), Confidence (ROUGE-L grounding), Novelty (1 - max cosine similarity), Recency, Type Prior. Learned weights vs fixed heuristics.
ACC: Agent Cognitive CompressorBousetouaneJan 2026Bio-inspired memory controller replacing transcript replay with bounded internal state updated online. Compresses agent cognitive state without losing decision-relevant context.
MemReader: From Passive to Active Extraction for Long-Term Agent MemoryKang, Li, Chen, Tang, Xiong, LiApr 2026First RL-trained active extraction policy (vs passive transcription). MemReader-4B uses GRPO + ReAct to evaluate value/ambiguity/completeness, then chooses WRITE / DEFER / RETRIEVE-CONTEXT / DISCARD before admission. Integrated into MemOS. Addresses memory pollution from noisy dialogue and cross-turn dependencies.
Personalize-then-Store: Benchmarking & Learning Personalized Memory for Long-horizon AgentsIn, Kim, Park, Yoon, Park (KAIST)May 2026Argues universal static storage policies waste budget on transient sessions while dropping critical long-horizon context. Introduces PerMemBench (first benchmark for personalized memory policies, with multi-year multi-domain personas) and session-level storage gating — a lightweight per-user admission filter that bypasses memory ops for transient sessions. Personalization yields large retention gains under perfect gating; accurate gating remains the open challenge 📄. ⚠️ Preprint.

Retrieval & Recall

PaperAuthorsDateKey Contribution
When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory (ReFind)Li, Zhang, Xu, Du, Fu, ChenAug 2026Asks how much of structured memory's benefit comes from the structure versus from competent retrieval over the raw history — and answers uncomfortably. ReFind builds no semantic structure at all: it leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls (session-aware rank fusion, local context expansion, temporal narrowing, skipping already-inspected sessions). Across ~2,800 questions under MemoryAgentBench's incremental multi-turn setting it attains the highest mean accuracy (58.2) of any system compared — above the strongest graph- and tree-based systems (HippoRAG 2, 53.2) — all under a GPT-4o-mini backbone matched to every reused baseline. Reaches 93.2±3.3 / 89.3±6.0 on LongMemEval-S/M with GPT-5-mini, with no LLM-based index construction at all 📄. ⚠️ Preprint.
xMemory: Beyond RAG for Agent MemoryHu, Zhu, Yan, He, GuiFeb 20264-level memory hierarchy. Key thesis: RAG ≠ agent memory. Submodular diversity-aware retrieval (MMR) replacing naive top-k. Uncertainty-gated adaptive expansion. Theme clustering. 44.91% retroactive reassignment rate proves flat structures fail. ⚠️ Conference status unverified.
SuperLocalMemory V3—Mar 2026First information-geometric foundations for agent memory. Fisher information metric replaces cosine similarity. Riemannian Langevin dynamics for retrieval. Theoretical grounding for memory operations.
ExpRAG: Retrieval-Augmented LLM Agents Learning to Learn from ExperienceFerraz, Deffayet, Nikoulina, Déjean, ClinchantMar 2026Trains agents to USE retrieved trajectories in-context (retrieval-augmented fine-tuning). Standard LoRA collapses on OOD tasks; ExpRAG-LoRA generalizes to held-out hard tasks. Combines experience retrieval with fine-tuning — neither alone is sufficient.
To Know is to Construct: Schema-Constrained Generation for Agent Memory (SCG-MEM)Zheng, Song, Li, YangApr 2026Replaces dense retrieval with schema-constrained decoding. Piaget-inspired: assimilation (grounding into existing schemas) vs accommodation (expanding schemas with novel concepts). Solves "structural hallucination" — LLMs generating references to non-existent memory keys. Constructivist alternative to similarity-based recall.
SAM: State-Adaptive Memory for Long-Horizon Agents—May 2026Targets retrieval of scattered information over long horizons — adapting memory access to evolving task state rather than fixed top-k recall. ⚠️ Preprint; venue unconfirmed.
RaMem: Contextual Reinstatement for Long-term Agentic MemoryYang, Kan, Li, Li, Qin, Li, Bogdan, ThomasonJun 2026Names and addresses context collapse: retrieved memory fragments that share entities or user states can look equally relevant even when their surrounding episodic conditions (time, session, participants) differ, so similarity-based retrieval returns content-relevant but context-invalid evidence. RaMem runs four coordinated stages — evidence anchoring, recall-condition induction, validity-aware retrieval, context-preserved synthesis — to prioritize context-compatible memories over merely similar ones. +10% average F1 over strong baselines across several backbones on long-term memory benchmarks 📄. ⚠️ Preprint.

Forgetting & Consolidation

PaperAuthorsDateKey Contribution
TEPA: Revoking Stale Memories for Conflict-Robust Language AgentsZhou, Ouyang, Zheng, XiangAug 2026Makes validity an explicit state of memory. Observations are stored as keyed precedents; an active precedent is revoked when fresh evidence contradicts it under the same key, so retrieval draws only from current evidence while revoked history is preserved for audit rather than deleted. The headline result is a failure mode for the naive strategies: under full reversal, append-only and last-write-wins both score 0.210 — below the 0.309 of having no memory at all — while TEPA reaches 0.950 (50 seeds; reproduced under real file-backed execution at 0.203 / 0.298 / 0.950). Honest scoping: on clean MemoryAgentBench SH-6k, TEPA only matches a strong last-write-wins cache, confirming current-key replacement as the decisive op for single-hop facts; multi-hop and very-long-context settings expose retrieval-chain bottlenecks beyond fact-level validity tracking 📄. ⚠️ Preprint.
Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent MemoryKang, Liu, Kai, Liang, Tang et al.Jul 2026Reframes retention-vs-consolidation as a budget-dependent operator choice, not an architectural commitment. Decomposes each operator's utility (Merge / Abstract / Rewrite) into a coverage effect on evidence that retention omits plus a signed replacement effect on raw evidence that already fits — their balance explains why the preferred action changes with relative budget pressure. Implemented as Offline Abstraction-Safety (OAS), a lightweight learner estimating action utility from pre-generation features with held-out harm calibration. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets while retention is preferable under loose ones; LoCoMo replicates the crossover at a smaller budget, consistent with its shorter evidence. Cross-note abstraction and merging generally outperform local rewriting 📄. ⚠️ Preprint.
Control-Plane Placement Shapes ForgettingYangJun 202613-configuration architectural study of where an LLM should sit in the memory pipeline — recall plane (extensively benchmarked) vs. control plane that mutates via supersede/release/purge (largely untested). Finding: deterministic primitives handle lexical/temporal forgetting but fail canonicalization (5% identifier-obfuscation, 0% cross-lingual); inscribe-time LLM fixes canonicalization (100%) but can't do intent-aware deletion (0%); a mutation-time hook recovers intent-aware deletion (78–85%) and lifts nearly every category simultaneously (91.7–93.2% overall) at $0.17/385-case run. Releases ForgetEval (1000-case + 385-case adversarial suite, 10-annotator Fleiss' κ=0.958) plus a 6-method Adapter Protocol. Core claim: production failures are predominantly forgetting failures, not recall failures — yet benchmarks measure only recall. Code/benchmark released (MIT) 📄.
SleepGate: Sleep-Inspired Forgetting for LLMsXie (Kennesaw State)Mar 2026Learned sleep cycle over KV cache addressing proactive interference. Conflict-aware temporal tagger + forgetting gate + consolidation module. Reduces interference horizon from O(n) to O(log n). 99.5% retrieval accuracy at PI depth 5 vs 23% baseline 📄 (self-reported; extraordinary gap warrants independent replication). Supersession detection via binary σ flag.
SCM: Sleep-Consolidated Memory with Algorithmic ForgettingShindeApr 2026Five components inspired by human memory: limited-capacity working memory, multi-dimensional importance tagging, offline sleep-stage consolidation with distinct NREM and REM phases, intentional value-based forgetting, and a computational self-model for introspection. Reports perfect 10-turn recall, 90.9% noise reduction, sub-millisecond search 📄. Research preview.
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent MemoryGu, Xiong, Wang, Ren, Li, Zhang, Guo et al.Apr 2026Grounds selective forgetting in hippocampal indexing/consolidation theory and the Ebbinghaus curve. Argues that in resource-constrained environments a forgetting mechanism is as crucial as retention — and that memory security depends on the agent's ability to actively drop sensitive history.
Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning (LRE)Lia, MazumderJun 2026Reframes eviction as a fidelity problem, not a compression problem: dropping one load-bearing detail (an access token, a required path) fails the task outright. Learned Relevance Eviction (LRE) is a few-kilobyte, CPU-only, LLM-free scorer that learns which history units are load-bearing and keeps them verbatim. Matches full-history-retention accuracy while cutting peak context up to 52%; on LoCoMo gives the best budgeted answer quality reading 68% fewer tokens; annotation-free training on the system's own behavior recovers 95% of supervised effectiveness. Evidence that cheap learned relevance can replace LLM-mediated summarization for eviction 📄. ⚠️ Preprint.

RL-Based Memory Policy

PaperAuthorsDateKey Contribution
AgeMem: Agentic Memory — Learning Unified LTM and STM ManagementYu, Yao, Xie, Tan, Feng, Li, WuJan 2026Exposes store/retrieve/update/summarize/discard as tool-based actions. Agent learns WHEN and WHAT to do via 3-stage progressive RL + step-wise GRPO. ACL 2026 SAC Highlight ✅ — Outperforms all heuristic-based baselines across 5 benchmarks. Key thesis: memory operations should be learned, not hard-coded 📄.
EMPO²: Exploratory Memory-Augmented On- and Off-Policy OptimizationLiu, Kim, Luo, Li, YangFeb 2026Uses memory not just for recall but for exploring novel states. Hybrid on/off-policy RL — 128.6% improvement over GRPO on ScienceWorld 📄. Adapts to new tasks with just a few memory-augmented trials, zero parameter updates. ICLR 2026 ✅.
MemFactory: Unified Inference & Training Framework for Agent MemoryGuo, Li, Tang, Xiong, LiMar 2026"LLaMA-Factory for memory agents" — first unified modular framework. Lego-like plug-and-play memory components. Natively integrates GRPO for RL-based memory policy training. Supports Memory-R1, RMM, MemAgent paradigms out of the box. Up to 14.8% improvement over base models 📄.
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM AgentsYan, Bahloul, Nie, Schwarzmann, Trivisonno, Tresp, MaMay 2026Identifies a fundamental flaw in GRPO for memory-RL: once rollouts write different memories, they no longer share the same effective environment, so trajectory-level group comparisons are unfair. Introduces LoGo-GRPO (Local + Global) — global keeps end-to-end long-horizon reward; local re-rollouts compare memory-op outcomes from the same intermediate state. Shared-parameter co-learning for fact extractor + memory manager. Progressive curriculum 8→16→32 sessions.
What Training Data Teaches RL Memory Agents: Curriculum Effects in Memory-Augmented QAHe, Lin, Liu, Wu, Xie, Zhou, XiaoMay 2026Controlled study holding architecture/RL/hyperparams fixed and varying only curriculum across in-domain (LoCoMo), mixed (LoCoMo+LongMemEval), and OOD (LongMemEval). Key findings: curriculum is a fine-grained lever on specialization, not a uniform scaling factor; per-type differences dwarf aggregate differences (single-number benchmark comparisons systematically underreport); binary EM reward produces no signal at G=4 group size — continuous rewards needed for single-GPU regime. Practical RL-memory training playbook. Code
SaliMory: Orchestrating Cognitive Memory for Conversational AgentsZhang, Zhang, Jiang, …, X. L. Dong (Meta)Jun 2026Trains a single LM to manage a cognitively-structured memory (user facts, preferences, working memory). Introduces a hierarchical stage-wise process reward + reward-decomposed contrastive (GRPO-style) refinement, giving isolated supervision to distinct ops (selective filtering, consolidation, cue-driven recall) end-to-end — sidestepping the credit-assignment bottleneck of standard RL over multi-stage pipelines. Cuts memory-attributed failures by ~⅓, +10% end-to-end accuracy, more than doubles the Good-Personalization rate 📄. ⚠️ Preprint.
JAMEL: Joint Agent Memory & Exploration Learning via Novelty SignalsTian, Weng, Kong, …, Y. Li (Tsinghua / AIR)Jun 2026Co-trains the memory module and exploration policy in a mutually-dependent loop: exploration needs memory to tell exhausted from unseen behaviors, while novelty-seeking interaction supplies annotation-free supervision (e.g., code coverage) for the latent memory. Generalizes to unseen environments; rivals a closed-source model while reducing tokens 📄. Code/model open-sourced. ⚠️ Preprint.

Context Management

PaperAuthorsDateKey Contribution
The Missing Memory Hierarchy: Demand Paging for LLM Context WindowsMason (UBC / Georgia Tech)Sep 2025Maps OS virtual memory concepts to LLM context: physical memory = context window, virtual memory = persistent state, page table = retrieval handles, page fault = re-request evicted content. Analyzed 857 sessions, 54,170 API calls, 4.45B effective input tokens.
What to Keep, What to Forget: A Rate–Distortion View of Memory Compaction in LLMs and AgentsColaco, LahjoujiJul 2026Unifies four separate compaction literatures — KV-cache eviction/quantization, prompt pruning/distillation, bounded architectural state, and agent memory consolidation — as instances of one rate–distortion problem: what to retain vs discard, at what fidelity, under a resource budget, to preserve downstream utility. Builds a layer-agnostic lower bound and a seven-axis taxonomy that lets mechanisms transfer between layers that have never been connected (serving-stack KV management ↔ agent long-term memory). Finding: at every layer the keep/discard signal is attention magnitude or recency, and it fails the same way everywhere — discarding before the query is known, with no way to undo it. Proposes a cross-layer benchmark that no existing benchmark provides.

Evaluation & Benchmarks

PaperAuthorsDateKey Contribution
Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory AgentsHuang, Zhang, Wu, Chen, Jiang, Yang et al. — incl. Ying Nian Wu, Kai-Wei Chang, Philip S. Yu, Aylin Caliskan (15 authors)Aug 2026Controlled harness comparing memory substrates — the underlying medium memory is represented in: dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms — across 3 backbones × 4 benchmark suites under 26 performance and efficiency metrics. Headline: no single substrate consistently dominates. Broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context; substrates that do well at moderate history lengths become costly or brittle at longer horizons. Motivates substrate routing as a necessary component of adaptive agent memory. ⚠️ Preprint; code on acceptance.
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent MemoryLi, Du, Xu, MaoJul 2026Names an assumption so natural it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks it — a tree-nut allergy should change the answer to a macaron request through almond flour, yet the two texts share no cue a retriever can see. InMind is a 125-task, expert-verified benchmark across ten life domains (113 tasks grounded in citable public sources) whose paired controls separate three explanations existing evaluations conflate: never stored, no bridging knowledge, or stored-but-never-surfaced. The verdict is clean: with the decisive memory placed in context the backbone answers 84.0% of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic systems reach at most 14.4% — even though they recall those same facts on demand at up to 100%. An embedding with 8× the dimensionality does not close the gap. Locates the failure in the query-conditioned interface itself, naming routing — deciding which facts stay visible — as the open problem.
MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory EvaluationWangJun 2026Controlled one-variable-at-a-time protocol on LongMemEval-S across 3 model families, exposing confounds that inflate architecture claims: verbatim RAG ≈ full-context GPT-4o-mini (47.2% vs 49.8%, p=0.34) but the ranking reverses by model (Gemini +14pp from full context; Sonnet +31pp from RAG, partly by refusing 63% of full-context queries); swapping only the embedding model shifts accuracy +6.2pp and flips which system wins; agent self-memory (42%) underperforms basic retrieval (47%); Mem0 matches cloud RAG on only 2/6 question types at 50x the cost. Recommends fixing embedding models across comparisons and reporting write-path cost before attributing gains to architecture — directly relevant to how Infini Memory/MAB-style comparisons should be read 📄.
StructMemEval: Evaluating Memory Structure in LLM AgentsShutova, Olenina, Vinogradov, SinitsinFeb 2026Tests agents' ability to organize memory (not just retrieve). Tasks: transaction ledgers, to-do lists, trees. Key finding: LLMs don't spontaneously recognize when to apply memory structure — they succeed only when explicitly prompted.
MemoryAgentBench: Evaluating Memory via Incremental Multi-Turn InteractionsHu, Wang, McAuley (UC San Diego)Jul 2025 (revised Jun 2026)Tests 4 memory competencies in realistic incremental accumulation. Key finding: no current method masters all 4 competencies simultaneously. ICLR 2026 ✅. Code
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (Memora + FAMA)Uddin, Shubham, Blanco, Baral, Wang (ASU)Apr 2026ACL 2026 Findings ✅ — Introduces Memora, a weeks-to-months benchmark over three tasks: remembering, reasoning, recommending. Introduces FAMA (Forgetting-Aware Memory Accuracy) — a metric that penalizes reliance on obsolete or invalidated memory rather than just rewarding recall. Evaluation of 4 LLMs and 6 memory agents finds frequent reuse of invalid memories and failures to reconcile evolving knowledge. ⚠️ Not to be confused with Microsoft's Memora: Harmonic Memory Representation (2602.03315).
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?Chao, Bai, Sheng, Li, SunMay 2026Benchmark for Implicit Conflict — a later observation invalidates an earlier memory without explicit negation, requiring inference + commonsense to detect. 400 expert-validated scenarios, 1,200 queries, contexts up to 150K tokens. Three-dimensional probing: State Resolution, Premise Resistance, Implicit Policy Adaptation. Best frontier LLM only 55.2% overall. Companion prototype CUPMem uses structured state consolidation + propagation-aware search. Complements FAMA.
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic TasksHe, Wang, Zhi, Hu, Chen, Yin, Wu, Ouyang, Wang, Pei, McAuley, Choi, PentlandFeb 2026Memory-Agent-Environment loop benchmark across web navigation, preference-constrained planning, progressive information search, sequential formal reasoning. Key result: agents near-saturated on LoCoMo perform poorly in MemoryArena, exposing a gap between long-context memorization benchmarks and interdependent agentic memory use. Project
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State RecoveryMa, Zhou, Huang, Yang, Ma, Wang, Li, Miao, Yu, WangJun 2026Argues memory should be evaluated as an auditable post-interaction artifact, not just through downstream task success. Simulates 50 users each carrying a hidden, taxonomy-anchored 31-dimension state bank across leak-controlled tasks, then reconstructs that state from the agent's resulting memory (full-store and top-k access) and scores it against ground truth. Key finding: task completion nearly saturates even for a memoryless baseline, while category-balanced state recovery stays moderate (~0.6) and drops further under top-k — successful assistance and recoverable memory are distinct capabilities. First benchmark to score memory recovery directly.
Are We Ready For An Agent-Native Memory System?Zhou, Zhou, Han, Xu, Li, Li, Xiong, WuJun 2026Systematic data-management-perspective evaluation decomposing agent memory into four modules — representation/storage, extraction, retrieval/routing, maintenance — and benchmarking 12 representative memory systems across 5 workloads / 11 datasets. Key finding: no single architecture dominates; effectiveness depends on whether the memory structure matches the workload bottleneck. Fine-grained ablations quantify effects on representation fidelity, retrieval precision, update correctness, long-horizon stability; localized maintenance beats global reorganization on cost. Code and paper-list released.

Cognitive & Neuroscience-Inspired

PaperAuthorsDateKey Contribution
Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory SystemsBhandari, Wadhwani, Kumar, NarangAug 2026Takes one specific mechanism from western scrub jay episodic memory — per-memory, type-conditioned temporal decay — and operationalizes it as an auto-classified perishability coefficient in an external store, against the default that all memories are equally persistent. Each memory is a jointly-bound What–Where–When tuple carrying estimated perishability and a utility horizon, retrieved by query-adaptive scoring and revised retroactively at O(1) LLM calls per update. Introduces the Temporal Generalization Test (held-out retention intervals) and a Generalization Gap metric, on which ScrubJay-MEM is the only retrieval-based system with substantially positive GenGap (+0.108); on MemoryAgentBench EventQA-64k it improves F1 by +2.66 over Mem0 and +3.09 over Qwen3-Embedding-4B, and a decay ablation collapses GenGap by 5.7×. Unusually candid scoping: gains narrow under stronger backbones and reverse on fact-consolidation tasks 📄. ⚠️ Preprint.
Episodic Memory is the Missing Piece for Long-Term LLM AgentsPink et al. (UT Austin)Feb 2025Position paper arguing episodic memory — supporting single-shot learning of instance-specific contexts — is the critical missing capability. Defines 5 properties: temporal, instance-specific, single-shot, inspectable, compositional.
MAP: Modular Agentic PlannerCorrea, Samwick, Gershman et al. (Microsoft Research, Harvard)Oct 2023 / Nature Comms 2025Brain-inspired architecture decomposing planning into PFC-associated modules: conflict monitoring, state prediction, state evaluation, task decomposition, task coordination. arXiv:2310.00194.
SCL: Structured Cognitive Loop with Governance LayerKimNov 2025R-CCAM: Retrieval → Cognition → Control → Action → Memory. Soft Symbolic Control = governance layer applying symbolic constraints to probabilistic inference. Zero policy violations, complete decision traceability.
Procedural Memory Is Not All You NeedWheeler, JeunenMay 2025LLMs are constrained by reliance on procedural memory (pattern-driven tasks). For "wicked" learning environments with shifting rules and ambiguous feedback, different memory types are required. ACM UMAP '25.
Evaluating Theory of Mind in Multi-Agent LLM SystemsKostka, Chudziak (Warsaw UT)Sep 2025ToM and Internal Belief mechanisms are NOT universally beneficial — stronger models handle extra cognitive load well, but weaker models get confused. Model capability is the dominant factor. ICCCI 2025 ✅ (journal_ref).

Skill & Procedural Memory

PaperAuthorsDateKey Contribution
Demystifying Agent Skills: Why They Work—Until They Don'tJiang, Huang, Xing, Wu, Gao et al.Aug 2026Moves skill evaluation past aggregate success rates by isolating representation, outcome annotation, retrieval difficulty, and cross-framework robustness. Normalizes 8,135 trial records into a taxonomy of three categories and twelve skill-use modes. Core mechanism finding: skills work because noisy trajectories become procedural anchors that stabilize execution — 65.7% of cases, versus 4.5% for explicit knowledge injection (+6.06 points over Workflow Memory in matched comparisons); skills stabilize action, they do not inject missing facts. The scaling warning is the more useful half: retrieval is a separate bottleneck — as pools grow from 5 to 100, actual-use precision falls 29.6% → 3.3%. Exact ground-truth invocation is found to be neither sufficient nor necessary. Directly relevant to anyone scaling a skill/procedure catalog 📄. ⚠️ Preprint.
Muscle Memory for Agents: Compile not Merely RetrieveOmran, Lanka, Zhang, DixitAug 2026Position paper against the field's default pattern (store experience → retrieve at inference → let a general-purpose orchestrator interpret it), arguing it is the wrong default for personalization. Proposes compiling recurring user intent into purpose-built specialist agents as a memory paradigm distinct from retrieval, targeting the multi-turn tax where users repeatedly correct format, depth, and scope. Reference implementation is a four-phase pipeline (Harvest → Analyze → Augment → Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable specialists with two-stage trigger matching. On 90 held-out scenarios across five personas it wins 32 of 36 cases where a specialist fires (88.9%), at +2.05 personalization gain and a −0.28 accuracy cost on a 1–4 scale — note the win rate is conditional on a specialist firing 📄. ⚠️ Preprint.
SkillRouter: Skill-Based Routing for LLM AgentsAlibabaSep 2025Critical finding: skill BODY is the decisive routing signal (91.7% attention weight), NOT name (7.3%) or description (1.0%). Removing body causes 29-44pp degradation. BM25 on metadata alone scores 0%. Two-stage retrieve-and-rerank (1.2B params).
EvoSkill: Self-Evolving Skill DiscoverySentient, Virginia TechSep 2025Automates skill discovery via iterative failure analysis without model fine-tuning. Three agents: Executor, Proposer, Skill-Builder. Skill-merge outperforms single runs. Code
Trajectory-Informed Memory Generation for Self-Improving Agent SystemsFang, Isahagian, Jayaram, Kumar, Muthusamy, Oum, Thomas (IBM Research)Mar 2026Extracts typed actionable tips from execution trajectories: strategy tips (clean successes), recovery tips (failure-then-recovery), optimization tips (inefficient-but-successful). Closes the gap between episodic logs and reusable procedural knowledge — useful as a complement to EvoSkill and SkillRouter.
DeltaMem: Incremental Experience Memory for LLM Agents via Residual TreesTan, Zhang, Cao, Li, Chen (RUC)Jun 2026Stores experience as residual deltas across two trees — goal-conditioned task skills and scene-level environment knowledge — so related episodes share a common root instead of duplicating content. Retrieval: failure-penalized similarity scan + root-to-match chain reconstruction; autonomous consolidation distills high-frequency paths into new roots. Cuts redundancy and retrieval conflicts; beats baselines across interactive environments 📄. Code released.

Multi-Agent Memory

PaperAuthorsDateKey Contribution
MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent SystemsXiong, Zhou, Wang, Gu, Tang, Li et al. (9 authors)Aug 2026Targets a path defenses structurally miss: a poisoned memory is written once, retrieved repeatedly, promoted into shared memory, and reused by other agents — so a single write steers many later decisions and contaminates agents that never saw the original attack, while no malicious message ever crosses a visible communication edge at the moment of harm. Because existing safeguards inspect prompts, actions, or communication edges, they miss content that looks benign at write time but becomes harmful after retrieval. MAPLE-Guard monitors the memory lifecycle and places gates at write, retrieval, promotion, and cross-agent reuse, quarantining risky memories and blocking poisoned private memories before they enter shared memory. ASR 38.2% → 0.9% (LongMemEval) and 34.7% → 0.2% (AppWorld); multi-agent defense success rate 54.0%→74.3% and 42.5%→99.8% 📄. Code ⚠️ Preprint.
Governed Shared Memory for Multi-Agent LLM Systems (MemClaw / ArgusFleet)Margalit, Cohen-Inger, Avram, Taig, MargalitJun 2026Formalizes the fleet-memory problem: unauthorized leakage, stale propagation, contradiction persistence, provenance collapse. Defines systems-level primitives — scoped retrieval, temporal supersession, provenance tracking, policy-governed propagation — implemented in MemClaw, a production multi-tenant memory service, evaluated live via ArgusFleet rather than a baseline comparison. Results: 100% correct depth-4 provenance reconstruction at sub-second/hop; zero cross-fleet leakage. Also reports real production bugs found via live eval: sub-tenant scope bypass on direct GET-by-id (disclosed/remediated) and a pipeline-ordering conflict where a synchronous near-duplicate gate can reject contradictory writes before the async contradiction detector runs. Conclusion: long-context retrieval alone is insufficient for production multi-agent memory 📄.
CoMAM: Collaborative Memory for Multi-Agent Systems—Sep 2025Collaborative RL framework modeling memory agents as sequential MDP with inter-agent dependencies. Group-level ranking consistency for coordinated memory operations.
DCM-Agent: Dual-Cluster Memory for Multi-Paradigm AmbiguityZhang, Wan, Zhang, Yang, Zhang, Wei, LiuApr 2026Tackles structural ambiguity in optimization problems where one problem admits multiple conflicting modeling paradigms. Training-free dual-cluster memory keeps competing paradigm solutions separated so the agent can pick or blend at inference time. Relevant to multi-agent settings where agents disagree on framing.

Memory Security & Privacy

April 2026 marks memory security emerging as a first-class research concern. As agent memories grow rich with user data, they become attack surfaces.

PaperAuthorsDateKey Contribution
SkillJack: Persistent Skill Backdoors in Self-Evolving AgentsYing, Wu, Wu, Zheng, Cheng et al.Aug 2026Prior poisoning work only bites when a poisoned record is retrieved; SkillJack attacks the experience-to-skill pipeline instead, hijacking the agent's own learning process so poisoned experiences become durable behavioral artifacts. Three named properties: sanitization whitewashing (malicious intent obscured during skill extraction), cross-layer promotion (transient experience becomes persistent capability), and persistence isolation — 80.0% of skill-mediated attacks survive deletion of the original poisoned records, so source-record cleanup is not a remedy. On SkillX, safety detection drops from 98.5% on poisoned trajectories to 11.4% on the extracted skills, with attack success 56.2% (SkillX) / 89.2% (Anything2Skill) over 150 shared trajectories; some implanted skills unintentionally activate on benign queries. Motivates provenance-aware skill lifecycle protection. Code released (Tencent/AI-Infra-Guard) 📄. ⚠️ Preprint.
Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall (PPMF)Xu, Xiao, Shao, Liu, LiJul 2026The attack-side twin of authority collapse. Identifies memory provenance laundering: during LLM-based consolidation, an external untrusted observation is rewritten as apparent user history or workflow support — preserving the action trigger while erasing the low-trust source that should limit its authority. Prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy consolidation. PPMF is lightweight middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. Vulnerable consolidated memories reach up to 1.000 ASR; with provenance, confirmation, and risk labels intact, no evaluated unauthorized high-risk action passes the gate while benign actions remain executable 📄. ⚠️ Preprint — arXiv comment reads "EMNLP2026 submitted" (submitted, not accepted).
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and RepairChen, Xie, Fu, Zhou, Yu, XuanJul 2026Traces the same malicious semantics across persistence, downstream consequence, and selective repair — the repair axis prior poisoning benchmarks omit. 310 cases drawn from 48 realistic contexts (code/science, daily life, office work), each following a controlled Write–Execute–Forget protocol in an isolated runtime, with evidence-based adjudication across seven lifecycle checkpoints over a 24-configuration matrix (2 agent harnesses × 4 memory backends × 3 LLM backends). Across all 24: malicious memory persists in 84.2% of cases and the full Write–Execute chain succeeds in 50.3%; among successfully poisoned cases 59.6% complete the full Execute chain and 56.1% achieve selective repair 📄. ⚠️ Preprint.
Securing LLM-Agent Long-Term Memory Against Poisoning (TMA-NM)LouckJun 2026Proves content- and lineage-based memory-poisoning defenses are malleable — an attacker can launder untrusted origin through the agent's own summarization, a trusted-tool echo, or manufactured corroboration, making poisoned content look benign and flipping its derivation edge to "trusted." Formalizes the malleability problem for the write-retrieve-act pipeline and proves a machine-checked separation theorem: write-time origin binding is necessary, non-malleable origin-bound authority with Sybil-resistant corroboration-gated elevation is sufficient. TMA-NM hits 0% attack success (vs. up to 68% laundering ASR for existing defenses) across 8 frontier models at full legitimate utility. Releases benchmark, harness, and machine-checked TLA+ models 📄.
When Agents Remember Too Much: Memory Poisoning Attacks on LLM Agents (GhostWriter / AM-Sentry)Torres, Shrestha, MisraJul 2026Targets personal assistant agents — the convergence of conversational + action-planning agents handling sensitive data via untrusted sources. GhostWriter attacks in two phases (injection → activation) and achieves ~98% injection rate, ~60% activation rate against SOTA agents, exploiting the lack of security-focused memory governance. Proposes AM-Sentry (memory-saving policy + memory-retrieval screen) which dramatically cuts GhostWriter's success while preserving utility.
ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive QueryingLyu, He, Wang, Hu, Li, Chen, Li, ChenApr 2026First systematic privacy attack on agent memory. Combines data-distribution estimation of a victim agent's memory with entropy-guided adaptive querying to maximize leakage. Achieves up to 100% Attack Success Rate on extracting stored memories 📄 — substantially outperforming prior attacks. Establishes the need for privacy-preserving memory designs as a first-class research direction.
SSGM: Stability and Safety Governed Memory for LLM AgentsLam, Li, Zhang, ZhaoMar 2026 (rev May 2026)Conceptual governance framework that decouples memory evolution from execution by enforcing consistency verification, temporal decay modeling, and dynamic access control before any memory consolidation. Provides a taxonomy of memory corruption risks: topology-induced knowledge leakage, semantic drift via iterative summarization, and consolidation hazards. Complementary defensive framing to ADAM.

Memory Trust & Integrity

Wave 7 (June–July 2026) surfaces a distinct concern from privacy/extraction attacks: can you trust what memory itself contains and does to reasoning — corrupted consolidation, sycophantic over-reliance on stored user views, and stale/current facts silently coexisting in the same bank?

PaperAuthorsDateKey Contribution
When Memory Becomes Authority: Benchmarking Authority Collapse at the Consolidation Boundary (AuthMem-Bench)Zhan, Zhang, Guo, Zhao, LiuAug 2026Consolidation imposes an implicit authorization boundary — it decides whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. Names authority collapse: consolidation preserves the claim while erasing the source constraints governing its authorized use, so the memory implies greater authority than its source permits. AuthMem-Bench is a controlled paired benchmark holding the focal claim and downstream task fixed while varying only source authority. Across 7 consolidators × 7 backbones, collapse appears in 48 of 49 configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%; end-to-end, automatically predicted and persisted authority labels cut the observed rate from 16.9% to 0.0% with benign task success essentially unchanged 📄. Memory must preserve not only what was learned but the authority under which it may be reused. ⚠️ Preprint.
Beyond Similarity: Trustworthy Memory Search for Personal AI Agents (MemGate)Zhang, Chen, Ma, Hu, He, Zhang, Liu, Yang, Zhang, JiaJun 2026Names memory search itself as a trust boundary: a semantically-similar memory can still be contextually inappropriate, causing cross-domain leakage, sycophancy, tool-call drift, or memory-induced jailbreaks. Evaluates A-Mem, Mem0, MemOS, and OpenClaw (real-world persistent-state agent) and finds long-term memory behaves as a durable control channel, not just a utility layer. Proposes MemGate — a 9M-parameter, 35.1MB plug-in inserted between the vector store and backbone LLM that applies a query-conditioned neural gate, turning raw similarity search into task-conditioned memory admission, with no LLM modification or inference-time judge required. Reduces memory-induced threats across frameworks/backbones while preserving utility 📄.
TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term MemoryYang, Paul, Srinivasan, Kulkarni, ChappidiJun 2026Targets errors introduced by the write/revise/delete pipeline itself — omission, corruption, hallucinated content — which become persistent system-state failures once stored. A Memory Transition Verifier scores update transitions on coverage, preservation, faithfulness; preference pairs among candidate updates drive preference-guided RL to directly optimize consolidation behavior. SOTA on MemoryAgentBench, HaluMem, Mem-alpha; +12.14 F1 on HaluMem extraction; cuts omission/corruption/hallucination by 40.1% / 79.1% / 50.0% vs the strongest baseline per error type.
MemSyco-Bench: Benchmarking Sycophancy in Agent MemoryXiang, Chen, Tang, Wei, Ning, Lin, Zhang, SuJul 2026Identifies memory-induced sycophancy: agents over-align with a user's stored prior views at the cost of factual accuracy or objective reasoning. Existing memory benchmarks check whether memories are stored/retrieved/updated correctly but not how they bias downstream reasoning. Five tasks probe whether an agent can reject memory as evidence, respect its applicable scope, resolve memory-vs-objective-evidence conflicts, track updates, and still personalize appropriately. First benchmark to treat memory reliance itself as a failure mode.
A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent MemoryShi, Tang, TungJul 2026Names ghost memory: old, current, and transition facts coexist unmarked in the memory bank and get mixed during retrieval, misleading the answer model about what's true now. Argues memory should be evaluated at three decoupled levels — bank maintenance, retrieval, answer-time resolution — since final QA accuracy hides where the failure actually occurs. ATMA is a state-aware overlay that keeps superseded/transition records, builds evidence packets for the query's requested state view, and exposes current/historical/transition labels to QA. On LTP (new LoCoMo Temporal Plus benchmark), Graphiti+ATMA improves conflict accuracy +0.240 absolute; on LoCoMo, raises temporal F1 from 0.0295 to 0.1705.

Memory Transactions, Versioning & Recovery

Wave 9 (July–August 2026) surfaces the window's strongest convergent signal: in roughly four weeks, multiple independent groups imported database transaction semantics wholesale into agent memory. The shared premise — a memory write is not a belief commit — is forced by the fact that stored memory now drives irreversible external actions. Staged writes, snapshot isolation, versioned heads, natural-language rollback, and cascading repair of derived records are becoming memory-layer primitives rather than database trivia.

PaperAuthorsDateKey Contribution
MemTX: Transactional Belief Commit for Stateful Agent MemoryLi, Wang, Lu, Chen, Li, Song, Zheng, CaiJul 2026The cluster's strongest entry and the source of its thesis. In shared-memory multi-agent settings one agent's write becomes another's premise, and eventually a tool call with real side effects — yet current systems treat every accepted write as immediately actionable truth, so a polluted tool result or a teammate's half-finished note can silently drive an irreversible action. Each record carries evidence, permissions, provenance, and validity; writes are staged inside snapshot-isolated transactions admitted by a validate-and-commit pipeline; irreversible tool calls are gated on in-flight belief state; and retracting a belief triggers typed cascading repair of its derived records and tool side effects. Two invariants — action-safety gating and cascade-repair completeness — are machine-checked via property-based testing and bounded exhaustive enumeration of 5.5M protocol states, with zero violations. Across five backbones from three model families it leads all eight baselines with paired-McNemar significance on four and ties the strongest fifth, and is the only method with zero downstream harm on every backbone 📄. "Backbone capability does not substitute for commit discipline." ⚠️ Preprint.
MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State RecoveryCui, Tang, Yao, Meng, Ma, JiaJul 2026A governance layer outside the answer model, supplying the transaction boundary existing systems lack: Ordered PatchTest verifies whether an update is actually supported by its source, a Temporal Resolver selects the visible version when facts conflict, and a durable snapshot journal restores application-visible state after a fault. On an item-disjoint audit it accepts all 60 supported originals and rejects all 179 hard negatives; under persistent multi-key faults on LongMemEval-S and LoCoMo states it restores the complete declared active map without knowing the actual physical write set. Highest average F1 across all twelve answer-model configurations on MemoryAgentBench FactConsolidation, beating Dense by 17.06–24.07 points in five representative settings 📄. ⚠️ Preprint.
ChronoMem: Version Control and Semantic Rollback for LLM Agent MemorySu, Xu, Zuo, BertinoJul 2026Attacks forward-only evolution: memory systems continuously accumulate, consolidate, and overwrite with no principled mechanism to inspect, version, or revert — leaving agents brittle under corrections, concept drift, and corruption, particularly after they have already been exposed to subsequent information. ChronoMem commits whole-memory snapshots at every write, maintains structured version histories, and maps natural-language undo intents to concrete historical versions through hybrid lexical + semantic retrieval, rank fusion, and reranking. Introduces a post-exposure counterfactual protocol: can the agent answer queries and summarize history as if future updates had never occurred? Integrated into Google's production-ready, open-source Agent Development Kit; claims the first open-source system and benchmark for systematic semantic global memory rollback. ⚠️ Preprint.
TARL: Transaction-Aware Reliable Ledgers for Executable Memory ManagementXiao, Xu, Zhang, Chen, ShiAug 2026Retires the binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification — choices that may share a label while producing fundamentally different memory states. TARL maps each statement to one of five executable actions, identifying the affected memory, resolving its temporal scope, comparing source reliability, and updating accepted, pending, and rejected ledgers. Trained by comparing the memory states produced by alternative update operations. Ships TARL-Mem, a benchmark with fine-grained action labels and next-state targets; improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption 📄. ⚠️ Preprint.

Memory Economics

PaperAuthorsDateKey Contribution
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory SystemsPollertlam, KornsuwannawitAug 2026The cost-side counterpart to accuracy benchmarking: Mem0, Hindsight, and Mastra Observational Memory against two reference strategies (fixed-size rolling window, full-transcript resubmission), two backbones, conversations up to 400 turns, every cost measurement paired with accuracy on 665 LoCoMo questions. Three findings, all awkward for the field: (1) serving cost cannot be predicted from conversation length and message size alone — a regression that tracks the reference strategies closely misses the memory systems by 18–69%, cost being driven by internal memory behavior; (2) break-even against full transcript is highly sensitive to system and backbone, ranging from the first tens of turns for the cheapest to never within 400 turns for the most expensive; (3) no system wins on both axes — accuracy spans 21–54%, and backbone choice drives cost as much as the memory system does 📄. ⚠️ Preprint.
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI AgentsAlake, Bernardis, Cayet, Engel, Hilloulin, Hong, Hosler, Kavantzas, Kossyk, Le, Patra, Talamadupula, Venzin (Oracle)Jul 2026Database-native memory substrate built on Oracle Database, framed as a full lifecycle: ingestion → extraction → consolidation → retrieval → summarization → revision/removal, with a layered active-core / passive-store split and explicit scope control across users, agents, threads. Reports 93.8% LongMemEval accuracy using ~10.7x fewer tokens than flat-history baselines. Notable as a large enterprise-vendor entry (13 authors, Oracle) validating the memory-as-lifecycle framing that most of this list converges on independently 📄.
Beyond the Context Window: Cost-Performance AnalysisPollertlam, KornsuwannawitSep 2025Compares Mem0-style fact-based memory vs long-context LLMs on LongMemEval, LoCoMo, PersonaMemv2. Break-even: memory system becomes cheaper after ~10 interaction turns at 100K context. Long-context wins on factual recall but memory is competitive on reasoning.
Agent Memory: Characterization and System Implications of Stateful Long-Horizon WorkloadsOmri, Gan, Broveak, Geens, He, Pentland, Verhelst, Weissman, Tambe (Stanford / MIT / KU Leuven)Jun 2026The first systems-level characterization of agent memory: (1) a system-oriented taxonomy along four axes; (2) a phase-aware profiling harness attributing cost to construction, retrieval, and generation; (3) characterizes 10 representative systems across two benchmark suites, showing how design choices shift cost across write vs read paths; (4) derives 10 system recommendations (construction scheduling, capability floors, amortization via query volume, freshness-latency tradeoffs, fleet-scale management). The production "memory economics" lens. 📄 ⚠️ Preprint.

Neuromorphic & Bio-Inspired Memory

Most agent memory research ignores 50 years of neuroscience. This section bridges that gap — presenting the biological mechanisms that may underlie what LLM memory papers have independently re-discovered (forgetting, gating, dual-phase encoding), alongside the spiking neural network implementations that model them directly. Cross-references to LLM papers below are 🔬 editor's synthesis.

Paper / ProjectAuthorsDateKey Contribution
A Bio-realistic Synthetic Hippocampus for Robotic CognitionTalanov et al.Oct 2025Synthetic hippocampal architecture with dual-phase operation: online sensorimotor encoding during wake, offline consolidation via SWR-triggered replay during sleep. SNN on neuromorphic substrates (≤1W). Goal-prioritised plasticity prevents catastrophic forgetting. Models the biological mechanisms that parallel what SleepGate and CraniMem implement in software 🔬. BioNanoScience. Open access.
The Memristive Implementation of the Hippocampus: A HypothesisTalanov et al.Aug 2025Hardware-level hippocampal memory using stochastic polycrystalline nano-fiber mesh as memristive substrate. Key insight: the inherent randomness of material structure mimics probabilistic biological synaptic networks. Demonstrates resistive switching tunability for implementing dynamic memory functions — bidirectional replay, synaptic up/down-scaling, consolidation. BioNanoScience. Open access.
Simulation of Serotonin Mechanisms in NEUCOGAR Cognitive ArchitectureTalanov, Gafarov, Vallverdú et al.2018Maps neuromodulatory mechanisms to computational models: dopamine → attention, serotonin → inhibition. The "cube of emotions" model. Demonstrates that mammalian emotional-state control via monoamine neurotransmitters can be re-implemented computationally. Foundation for neuromodulator-gated memory admission. Procedia Computer Science.
tinyHippoTalanovActiveCA1 + CA3 hippocampal microcircuit simulation in NEST with Izhikevich neurons. Implements bidirectional replay, theta-modulated encoding/retrieval phase separation, and SWR-triggered consolidation. The biological reference implementation — validates the mechanisms that Membrain engineers and SleepGate abstracts. MIT license.
MembrainFatykhovActiveNeuromorphic memory bridge using FlyHash encoding and BiCameralMemory SNN (Nengo/Voja learning). Hopfield-style attractor dynamics for pattern completion (tested up to 20% noise in PoC). Stochastic consolidation with SleepSignal. gRPC API for agent integration. The engineering abstraction layer between biological models (tinyHippo) and cognitive agents (Nous).

The Stack: These projects form a natural hierarchy — tinyHippo (biological model, validates mechanisms) → Membrain (engineering abstraction, SNN service) → Nous (cognitive agent, consumes memory). The papers provide the theoretical foundation; the repos provide working implementations.

Why this matters for agent memory 🔬: The LLM papers in this list independently converge on mechanisms that neuroscience has studied for decades. The following correspondences are editor's synthesis — the LLM papers don't explicitly cite these neuroscience sources, but the structural parallels are striking:

  • SleepGate's learned forgetting ↔ hippocampal SWR consolidation during sleep (Talanov 2025a)
  • CraniMem's utility gating ↔ neuromodulator-gated admission (Talanov/NEUCOGAR 2018)
  • A-MEM's dual-phase encoding ↔ online/offline hippocampal states (Talanov 2025a, tinyHippo)
  • xMemory's hierarchical consolidation ↔ cortical-hippocampal memory transfer (Talanov 2025b)

🛠 Implementations & Reference Code

Papers describe mechanisms; these repos run them. Inclusion bar: (1) implements a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance; (2) actively maintained and (3) an OSI-approved license — both enforced for the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where a reader may actually vendor the code; author code for papers and benchmark/dataset repos is held to a different standard — reproduction value — so it may carry restrictive or absent license terms and may go quiet after publication, with both facts flagged inline ⚠️; (4) and either backs a paper in this list, has real production adoption, or fills a mechanism niche nothing else covers. Agent harnesses that merely bundle a vector store, general-purpose vector databases, and repos whose only claim is a self-reported benchmark ranking are deliberately excluded. Stars, license, and last-push verified via the GitHub API on 23 Aug 2026 — a snapshot for triage, not a ranking.

Memory Layers & Frameworks

ProjectStarsLicenseWhat it actually implements
mem063.8kApache-2.0Extraction-based fact memory: an LLM pass distills durable facts from raw dialogue, then reconciles them against the existing store via add / update / delete / noop. The most widely deployed reference for the extract-then-reconcile write path. Paper: Mem0.
Graphiti (Zep)30.2kApache-2.0Bi-temporal knowledge graph: edges carry validity intervals and are invalidated, not overwritten, so superseded facts remain queryable with time bounds. The closest production analogue to Memory Transactions, Versioning & Recovery.
cognee30.2kApache-2.0ECL (extract → cognify → load) pipeline turning documents and conversations into a typed knowledge graph plus vector index, with explicit prune/forget operations over the graph rather than append-only growth.
Letta24.4kApache-2.0The MemGPT lineage: OS-style paged memory where the agent edits its own core memory block through tool calls and pages the remainder to archival storage. Self-editing memory as the mechanism. Paper: MemGPT.
Hindsight20.9kMITConsolidation as an explicit policy layer over four levers — importance, merge, decay, eviction — instead of an unbounded append log. Industry framing already cited in Adjacent Research; paper: arXiv:2512.12818.
MemOS10.9kApache-2.0"Memory OS": a MemCube abstraction unifying plaintext, activation (KV), and parameter memory, with scheduling and cross-task reuse between them. Rare in treating KV-cache and long-term store as one substrate.
Honcho6.8kAGPL-3.0Memory as a representation of a person, not a transcript: background reasoning maintains per-peer models, queried through a dialectic API rather than similarity search.
MemMachine3.2kApache-2.0Clean two-tier service — episodic memory plus a durable user profile — with pluggable stores and a server/client split. A readable reference implementation if you are building your own.
MemoryOS1.6kApache-2.0EMNLP 2025 Oral ✅. Short / mid / long-term stores with heat-based promotion and eviction between tiers — one of the few systems whose admission and consolidation policy is inspectable code rather than a prompt.
redis/agent-memory-server308Apache-2.0Explicit working-memory (token-bounded, session-scoped) vs long-term memory split, with automatic extraction at the boundary. MCP + REST, Redis-native.

Paper → Code

Reference implementations for papers already indexed in this list — verified to exist and to correspond to the cited work.

Paper (in this list)CodeNotes
xMemoryHU-xiaobai/xMemory⭐120 · MIT. Author release for Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation.
AgeMemy1y5/AgeMem⭐37 · ACL 2026 SAC Highlight ✅. Unified long/short-term memory management learned as tool operations.
CraniMemPearlMody05/CraniMEM⭐7 · gated-and-bounded memory implementation; project page. ⚠️ No license file — all rights reserved by default; clear reuse with the authors.
HeLa-MemReinerBRO/HeLa-Mem⭐16 · ACL 2026 Main ✅. Hebbian associative distillation, reproduces the LongMemEval-S numbers. ⚠️ No license file — all rights reserved by default.
StructMemEvalyandex-research/StructMemEval⭐11 · Apache-2.0. Raw benchmark data + supplementary code from Yandex Research; authors label it work in progress.
memorywiremthamil107/memorywire⭐2 · Apache-2.0. Reference implementation of the wire format, including the provenance-based purge/quarantine path for poisoned stores.
A-MEMagiresearch/A-mem⭐1.2k · MIT. Zettelkasten-style memory that links and evolves notes on write. Paper reproduction lives in WujiangXu/A-mem (NeurIPS 2025 ✅). Last push Dec 2025 ⚠️.

Local-First & Coding-Agent Memory

Single-binary or file-backed systems, usually MCP servers, where the store is inspectable by a human and runs without a cloud dependency.

ProjectStarsLicenseWhat it actually implements
agentmemory27.3kApache-2.0Full lifecycle memory for coding agents: confidence scoring, decay and consolidation, knowledge graph, hybrid search, ~20 agent adapters. ⚠️ Headline ranking claims are self-reported.
Engram6.1kMITAgent-agnostic Go binary — SQLite + FTS5, MCP/HTTP/CLI/TUI, explicit conflict handling. Deterministic lexical recall as a deliberate counter-position to embedding everything (cf. ReFind in Retrieval & Recall).
memsearch (Zilliz)2.5kMITMarkdown files as the source of truth, Milvus for hybrid retrieval, typed episodic/procedural memory. Human-readable store = auditable memory.
Nocturne Memory1.3kMITRollbackable, visually inspectable graph memory server for MCP agents — edits are reversible operations, not blind writes. The practical counterpart to MemTX/ChronoMem-style recovery.
deja-vu1.0kMITNo write path: indexes the session transcripts 34 coding agents already write locally (lexical by default; optional deja embed may send text to an external embedding API) and serves them to any of the others — including pre-install history, so no cold start. Read-side mechanisms: outcome-aware ranking (a session that ended in a fix outranks one that gave up), redaction at ingest, per-project trust policy, forget/unforget. LoCoMo/LongMemEval drivers in-repo.
Vestige608AGPL-3.0Local-first Rust MCP server combining FSRS-6 decay, spreading activation, contradiction inspection, active suppression, Receipt Lock, and an inspectable dashboard. Built for causal recall — the quiet change that broke today, not the lookalike.
Mnemon507Apache-2.0Single Go binary, four-graph knowledge store with intent-aware recall, importance decay, and automatic deduplication. No API keys, one setup command.
Caura (ex-MemClaw)439Apache-2.0Governed shared memory for agent fleets: trust tiers, keystone policies, audit trails, multi-tenant scoping. The implementation counterpart to Memory Trust & Integrity and Multi-Agent Memory.
mengram189Apache-2.0Typed semantic / episodic / procedural memory where procedures are derived from recorded failures — one of the few OSS systems implementing the failure → procedure loop from Skill & Procedural Memory.
causal-memory66Apache-2.0Local-first Rust MCP server (also CLI + PyO3 bindings) where the store is one SQLite file holding facts, temporal state, and typed decision→outcome edges (caused / enabled / prevented). Spreading activation runs both ways — prevented edges carry negative activation, a GABA analogue — alongside Hebbian co-occurrence, Q-value dynamics, write-time gating, and immutable SWR-style consolidation. Reproducible benchmark harnesses (CausalEval, LoCoMo, LongMemEval) live in-repo; Hermes Agent provider plugin included.
Tree Ring Memory13MITProject-scoped Rust CLI lifecycle layer: SQLite/FTS explainable recall, deterministic consolidation, forgetting, redaction, audit, framework discovery. Memory that ages instead of accumulating raw transcript.
kgai1MITAppend-only, content-addressed decision log for dev teams: every change is a new event that supersedes the prior one through an explicit link with a reason, rejected paths stay queryable (temporal versioning), recall returns only decisions in force. Deterministic graph projection, lexical recall with no embeddings, and per-writer shard sync over the team's own S3 bucket (no server). Repo-supplied config is gated behind an explicit trust approval (governance). Claude Code plugin plus a Go CLI.
Hyperconsciousness1MITRust knowledge store with signed, encrypted append-only records and scoped, expiring grants for MCP access. Delegated grants cannot widen scope, add actions, or outlive their parent. Governance at the retrieval boundary; does not sandbox processes running as the owner. ⚠️ Developer alpha; no independent security audit.

Neuromorphic reference implementations — tinyHippo and Membrain — live in Neuromorphic & Bio-Inspired Memory.

Benchmarks & Evaluation Harnesses

ProjectStarsLicenseWhat it measures
LongMemEval1.0kMITICLR 2025 ✅. 500 questions across five long-term memory abilities over long interactive histories. The de-facto standard — and the number most vendor claims cite.
LongMemEval-V2132Apache-2.0Official successor, released 2026, addressing saturation and contamination in the original.
LoCoMo1.1ksee repo ⚠️Very long-term conversation benchmark (ACL 2024 ✅). Widely used dataset; frozen ⚠️ — last push Aug 2024, cite the data, don't expect maintenance.
MemoryData139none ⚠️Unified harness — 4 benchmark families, 22 method presets, one runtime — built specifically to make heterogeneous memory systems comparable. Directly attacks the fragmentation problem Evaluation & Benchmarks documents.
GateMem197MITMemory governance under multiple principals sharing one store: utility, access control, and active forgetting measured together instead of accuracy alone.
HaluMem155CC-BY-NC-ND-4.0 ⚠️Operation-level hallucination benchmark: scores extraction, updating, and answering separately, exposing errors that end-to-end QA accuracy hides. License is declared by README badge only — no LICENSE file in the repo; NC-ND terms forbid commercial use and derivatives, so treat it as read-only unless you clear it with the authors.
Mem-Gallery103MITACL 2026 Main ✅. Multimodal long-term conversational memory for MLLM agents. Last push Jan 2026 ⚠️.
STATE-Bench77MITMicrosoft. Memory-agnostic enterprise workflows measuring whether an agent improves with experience, not whether it recalls a fact. Blog write-up in Adjacent Research.
ResourceStarsWhy it's here
Agent Memory Techniques92630 runnable notebooks covering buffers, vector stores, KGs, episodic/semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, and LoCoMo evaluation. The best hands-on on-ramp before reading the papers above.
Awesome-AI-Memory (IAAR Shanghai)1.2kBroad bilingual knowledge base spanning LLM memory and agent memory, including engineering frameworks and applications.
Awesome-Agent-Memory (TeleAI)596Systems + benchmarks + papers for LLM and MLLM memory — stronger multimodal coverage than this list.
Awesome-GraphMemory339Survey-grade depth on graph-based agent memory specifically.
Awesome-Agent-Memory-Papers229Paper-only index with a browsable website.
OpenDataBox/awesome-agent-memory54Same name, different list — a four-axis paper taxonomy, companion to the MemoryData harness above.

Adjacent Research

These papers and projects aren't solely about agent memory but contribute relevant architectural ideas:

Paper / ProjectDateRelevance
MemTools: A Unified Research Framework for Interoperable Agent Memory — Zhao, Chen, Liang, He, Wang, Zhao, LiuJul 2026Targets architectural fragmentation as the thing blocking systematic memory research: implementations couple lifecycle stages together, entangle evaluation logic with specific datasets, and barely support heterogeneous memory types. MemTools standardizes the memory lifecycle through declarative data contracts, enabling interchangeable assembly of components across systems, and orthogonally separates benchmark datasets from execution protocols so evaluation can be reconfigured independently of data. Adds a unified computational interface for coordinating symbolic, neural, and multimodal representations in one runtime. ⚠️ Authors label it work in progress; the evaluation demonstrates isolation of design variables rather than task-level performance.
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories — Lovén, Sauvola, Riekki, TarkomaAug 2026Agents share transport and can call each other's tools, but have no protocol for reconciling what they know. MELD admits every incoming claim through a five-outcome procedure (insert, merge, relate, conflict, reject) decided from three signals — scoped claim-key identity, embedding similarity, and a natural-language- inference verdict — under context and freshness gates, acting through exactly one auditable, authenticated Patch as the sole object that mutates state, with a per-claim status CRDT over standard pub/sub. The design stance is the notable part: MELD does not adjudicate truth — a detected contradiction is preserved for later adjudication, never silently resolved. Merge classifier separates at AUC 0.968 with a 0.013 false-merge rate; the status CRDT reconverges in 30/30 real partition-heal trials where last-writer-wins manages 11/30; recall-non-inferior to a centralized store at ~11% less live storage; ~3× fewer messages at matched recall. Evaluated on a real continuum spanning operator-grade 5G edge, national HPC, and a local tier 📄. ⚠️ Preprint; code and data on Zenodo.
memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations — MunirathinamMay 2026 (rev Jul 2026)JSON-Schema 2020-12 wire format for 5 memory ops (remember/recall/forget/merge/expire) × 4 memory types over mem0/Letta/Cognee/Zep/pgvector backends, with an optional HITL governance channel. Not a new algorithm — a packaging of RRF/FSM/consolidation into a vendor-neutral protocol meant to compose with MCP. v3 adds a poison-recovery eval (PurgeBench) showing its provenance field is the strongest lever for recovering a poisoned store — relevant complement to TMA-NM/GhostWriter above.
EverMemOS — Liu, Bai, Chen et al.Jan 2026Self-Organizing Memory Operating System for structured long-horizon reasoning.
AriGraph — Anokhin, Radionov et al.IJCAI 2025 ✅Knowledge Graph World Models with episodic + semantic memory. Proceedings
Memory Matters More Than You Think — Zhou, Li et al. (Renmin U.)Jan 2026Event-centric memory with Logic Map for agent searching and reasoning.
The Consolidation Problem in Agent Memory — Vectorize (Hindsight)May 2026Industry framing of consolidation as a policy layer with four levers — importance, merge, decay, eviction — benchmarked across Mem0, Zep, Letta, LangChain, Hindsight.
Introducing STATE-Bench — Microsoft (repo)May 2026Open-source, memory-agnostic benchmark measuring whether agents improve with experience on stateful enterprise tasks (travel, support, shopping) — not just recall.
State of AI Agent Memory 2026 — Mem02026Industry survey of benchmarks, architectures, and six open problems (temporal abstraction, cross-session structure/identity, app-level eval, privacy/consent, staleness).

Key Themes & Cross-Cutting Insights

These papers, taken together, converge on several architectural principles:

1. 🚪 Gated Admission — Not Everything Deserves to be Remembered

Papers: A-MAC, CraniMem, SleepGate, ACC

The strongest systems filter before storing. CraniMem uses goal-conditioned cosine gating; A-MAC scores across 5 dimensions (utility, confidence, novelty, recency, type). The alternative — storing everything and retrieving selectively — leads to noise accumulation and proactive interference.

2. 📦 Typed Multi-Store — One Vector Store Is Not Memory

Papers: xMemory, SYNAPSE, AdaMem, CraniMem, Mem0, Episodic Memory

Every serious architecture separates memory by type (episodic, semantic, procedural, working). xMemory's 4-level hierarchy shows 44.91% of memories get retroactively reassigned — proving flat structures fundamentally fail. The human brain doesn't use one store; neither should your agent.

3. 🔄 Consolidation Loops — Memory Must Evolve

Papers: SleepGate, CraniMem, SYNAPSE, A-MEM

Raw episodic traces should consolidate into durable semantic knowledge over time. SleepGate's sleep cycles, CraniMem's replay-based promotion, and SYNAPSE's spreading activation all implement variants of this. Without consolidation, memory becomes a write-only log.

4. 🗑️ Forgetting as a Feature — Not a Bug

Papers: SleepGate, Procedural Memory, xMemory

SleepGate reduces proactive interference from O(n) to O(log n) through learned forgetting. This is the most under-implemented capability in agent frameworks — most systems only add memories, never remove or decay them. Forgetting is what keeps memory useful rather than just large.

5. 📊 RAG ≠ Agent Memory

Papers: xMemory, Memory Survey, Anatomy of Agentic Memory

The clearest thesis across this literature: retrieval-augmented generation is a retrieval strategy, not a memory system. Agent memory requires admission, evolution, consolidation, forgetting, typed storage, and context-aware retrieval. RAG addresses only the last item.

6. 🧪 Evaluation Is Broken

Papers: Anatomy of Agentic Memory, StructMemEval

Current benchmarks are saturated, metrics misalign with semantic utility, and results vary wildly by backbone model. StructMemEval reveals agents can't organize memory autonomously. The field needs better evaluation before it can claim progress.


7. 🎮 RL Is Replacing Heuristic Memory Management

Papers: AgeMem, EMPO², MemFactory, Memory-R1

The dominant shift in 2026: from hand-coded admission/retrieval rules to reinforcement-learned memory policies. AgeMem exposes all memory operations as tool-based actions and learns when to use them via GRPO. EMPO² takes this further by using memory for exploration, not just recall. MemFactory provides the infrastructure to train these policies. This is the most consequential paradigm shift since the field recognized that RAG ≠ memory.

8. 🔒 Memory Privacy Becomes a First-Class Concern

Papers: ADAM, FSFM

April 2026 surfaced the first systematic memory-extraction attack (ADAM, up to 100% ASR) — and the first forgetting frameworks (FSFM) that explicitly frame selective forgetting as a privacy primitive, not just a retention/efficiency one. Expect memory threat-modeling, audit interfaces, and unlearning APIs to follow.

9. 🪜 Memory ↔ Skills ↔ Rules as a Compression Spectrum

Papers: Experience Compression Spectrum, Externalization in LLM Agents

The memory and skill-discovery communities have been solving the same problem (extract reusable knowledge from interaction traces) without citing each other — cross-community citation rate is below 1%. The Compression Spectrum reframes both as points on a single axis (5–20× → 50–500× → 1000×+ compression), and identifies the missing diagonal: no current system supports adaptive cross-level compression.

10. 🧱 Active, Constructive Memory Operations

Papers: MemReader (active extraction), SCG-MEM (schema-constrained generation), HeLa-Mem (Hebbian distillation)

Two paradigm shifts in April 2026: (1) MemReader moves extraction from passive transcription to an RL-learned WRITE/DEFER/RETRIEVE-CONTEXT/DISCARD policy. (2) SCG-MEM replaces dense retrieval with schema-constrained decoding (Piaget-style assimilation vs accommodation). Together they suggest the future of memory is generative-constructive, not retrieve-and-paste.

11. 🧩 Reconstruction ≠ Retrieval — Memory Access as Active Reasoning

Papers: MRAgent, MemPro, DeltaMem

Wave 6's headline shift: stop treating memory access as a static retrieve-then-reason step. MRAgent folds LLM reasoning into graph traversal (Cue-Tag-Content + active reconstruction), iteratively pruning paths on accumulated evidence. MemPro generalizes the idea to the whole pipeline — the memory system itself becomes an evolvable program that rewrites its own construction/retrieval logic. DeltaMem reconstructs each experience on demand by composing a root-to-leaf residual chain rather than storing flat copies. Memory is increasingly constructed at access time, not fetched.

12. 📉 Memory Economics Gets Measured

Papers: Agent Memory Characterization, Beyond the Context Window, Personalize-then-Store

For two years memory papers reported accuracy; few reported cost. Wave 6 changes that. The first systems characterization (Omri et al.) builds a phase-aware cost profiler and a 4-axis taxonomy, deriving 10 deployment recommendations across the write and read paths. Personalize-then-Store reframes admission as a budget problem — gating transient sessions to spend the memory budget where it pays off. The question is shifting from "does memory help?" to "what does memory cost, and where does it amortize?"

13. 🔐 Trust & Integrity Become Explicit, Distinct from Privacy

Papers: TRUSTMEM, MemSyco-Bench, A-TMA

Wave 6 established memory as an attack surface (ADAM). Wave 7 shows a second, orthogonal failure class: memory that's technically retrieved correctly but still misleads. TRUSTMEM targets self-inflicted corruption during consolidation (omission, hallucination). MemSyco-Bench targets over-reliance — agents favoring a user's stored prior view over current evidence. A-TMA targets ghost memory — stale and current facts silently coexisting in the same bank. None of these are extraction attacks; they're reliability failures baked into ordinary write/retrieve operations.

14. 🔬 Evaluation Moves from "Did the Task Succeed" to "What Does the Memory Actually Contain"

Papers: MemProbe, Are We Ready For An Agent-Native Memory System?

MemProbe's headline finding — task completion nearly saturates even for a memoryless baseline, while structured state recovery from the memory artifact itself stays around 0.6 — is a direct rebuke of end-to-end accuracy as a memory metric. "Are We Ready" pushes the same point systemically: decomposing 12 memory systems into 4 modules shows no architecture dominates, and effectiveness is bottleneck-specific, not a single leaderboard number. Both extend Anatomy of Agentic Memory's Feb-2026 critique from theory into large-scale measurement.

15. 🧮 Compaction Gets a Unifying Theory

Papers: What to Keep What to Forget (Rate–Distortion View), Is Agent Memory a Database?

Two Wave 7 papers step back from architecture-of-the-week and ask for formal foundations. The rate–distortion paper shows KV-cache eviction, prompt pruning, bounded recurrent state, and agent memory consolidation are the same budget-constrained retain/discard decision, and that all four fail identically (irreversible, pre-query, recency/attention-biased). "Is Agent Memory a Database?" argues memory correctness is a property of the state trajectory, not individual records — no record-level storage system can satisfy it, motivating the GEM formalization. Both suggest the field's next gains come from theory, not another heuristic gate.

16. 🔑 Authority Is a Separate Property from Content — and Consolidation Destroys It

Papers: AuthMem-Bench, PPMF

Two independent papers landed the same finding within days, from opposite directions. Summarization is lossy in a way nobody was measuring: it preserves the claim and drops the permission. AuthMem-Bench measures it as an accident — authority collapse in 48 of 49 configurations, a 50.3% mean unauthorized-action rate for collapsed memories. PPMF weaponizes it — laundering an untrusted observation into apparent user history, preserving the trigger while erasing the low-trust source, up to 1.000 ASR. Both fixes are the same shape: persist authority as first-class metadata and match action risk against it at tool-call time. This is distinct from insight #13, which asks whether a memory is true; this asks what a memory is allowed to authorize.

17. 🧾 A Memory Write Is Not a Belief Commit

Papers: MemTX, MemTxn, ChronoMem, TARL

The transactions section in one line. Agent memory is acquiring database semantics — staged writes, snapshot isolation, versioned heads, rollback, and cascading repair of derived records and tool side effects. The forcing function is that memory now drives irreversible external actions, so the cost of an unsound write is no longer a bad answer but a real-world consequence. Note the direction of travel: TARL replaces a binary Write/Hold with five actions; ChronoMem adds an undo; MemTX gates tool calls on in-flight belief state. All three treat the moment of writing as provisional by default.

18. 🧱 Structure Is Not Free — and May Not Be Necessary

Papers: ReFind, Filesystem-Based Memory, Harness the Memory, Keep It InMind

The sharpest contrarian thread this list has carried. ReFind beats HippoRAG 2 (58.2 vs 53.2) with no semantic structure at all, on a matched backbone. Filesystem-Based Memory finds organization buys retrieval economy, not accuracy — and erodes as stores grow. Harness the Memory finds no substrate dominates and that excessive retrieval actively harms sequential decision-making. Against all three, Keep It InMind argues the framing itself is wrong: the failure isn't structure versus no-structure but the query-conditioned retrieval interface — 84.0% with memory in context collapsing to ≤14.4% when the same memory must be retrieved. Read together they pose an uncomfortable question for most of this list: is the field over-investing in write-time structure to compensate for a broken read-time interface?

Open Questions

  • How should admission gates interact with consolidation loops? (A-MAC + SleepGate integration)
  • What's the right forgetting curve for different memory types? (No paper addresses type-specific decay)
  • Can skill routing be unified with memory retrieval? (SkillRouter + xMemory convergence)
  • What's the minimum viable memory architecture? (Most papers add complexity — none simplify)
  • How should backward verification loops (MemMA probe-verify-repair) integrate with forward memory policy? (MemMA + AgeMem convergence)
  • Can RL-learned memory policies transfer across domains? (EMPO² shows OOD promise — but limited evaluation)
  • Can constructivist memory (schema-bound decoding, SCG-MEM) replace similarity-based retrieval?
  • Can the "missing diagonal" be filled — adaptive cross-level compression between episodic/procedural/declarative?
  • How robust is agent memory against systematic extraction attacks (ADAM-class threats)?
  • What is the right interface between active extraction (MemReader) and admission gating (A-MAC)?
  • Does Hebbian distillation (HeLa-Mem) outperform LLM-driven semantic abstraction at scale?
  • Can access-time reconstruction (MRAgent) scale without latency blowup versus precomputed retrieval?
  • What is the right cost model for agent memory at fleet scale, and when does a memory system amortize its write cost? (Agent Memory Characterization)
  • Do trust verifiers (TRUSTMEM-style) generalize across consolidation architectures, or must every memory system train its own?
  • Is memory-induced sycophancy (MemSyco-Bench) an architectural property or a training-data artifact inherited from the base LLM's own sycophancy?
  • Does ghost memory (A-TMA) require an architectural fix, or is decoupled bank/retrieval/answer-level evaluation enough to catch it before deployment?
  • Can a single rate–distortion budget metric make KV-cache eviction, prompt compression, and agent memory consolidation apples-to-apples comparable?
  • If task success saturates even for memoryless baselines (MemProbe), how many published memory-system benchmark wins are actually measuring memory at all?
  • If an agent with controllable lexical search over raw logs (ReFind) matches graph memory, which workloads actually justify the cost of write-time structure?
  • Can authority metadata (AuthMem-Bench, PPMF) survive consolidation without a full transactional substrate (MemTX), or are the two the same requirement discovered from different ends?
  • What is the right granularity for memory rollback — whole-store snapshots (ChronoMem) or per-record validity intervals (MemTX, TARL)?
  • Does native in-backbone memory (Metis) subsume external memory systems, or does it just relocate the same admission, forgetting, and provenance problems inside the weights where they can no longer be audited?
  • If retrieval fails on implicit associations (Keep It InMind) even at 8× embedding dimensionality, is routing-by-visibility a fix or an admission that the retrieve-on-query interface is the wrong abstraction?

Contributing

This list grows as the field grows. To contribute:

  1. Open an issue or PR with the paper title, arXiv/DOI link, authors, date, and a 1-2 sentence summary of the key contribution
  2. Suggest which section it belongs in (or propose a new one)
  3. Bonus: note how it relates to or challenges existing papers in the list

Papers should be peer-reviewed, at reputable workshops (e.g., ICLR MemAgents), or highly cited preprints. We prioritize papers with novel architectural ideas over incremental benchmark improvements.

For repositories (the Implementations section), the bar is different: a repo must implement a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance — and either back a paper in this list, show real adoption, or fill a niche nothing else covers. Star counts are context, not a criterion. Self-reported benchmark rankings are not evidence; a reproducible harness is. Active maintenance and an OSI-approved license are required only in the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where the code is meant to be vendored. Author code for papers and benchmark/dataset repos is judged on reproduction value instead: a license must still be disclosed, and where it is absent or non-OSI (e.g. NC-ND, or no LICENSE file at all) — or where the repo has been quiet for more than six months — the row is flagged ⚠️ rather than dropped. A frozen benchmark that everyone cites still earns its place; it just gets its last-push date stated.


License

CC0

This list is released under CC0. The papers themselves retain their original licenses.

agentic-ai
agent-memory
awesome-list
cognitive-architecture
llm-agents
memory-systems
neuromorphic