A curated collection of research papers on memory systems for LLM-based agents — covering architecture, retrieval, forgetting, consolidation, evaluation, the cognitive science that inspires it all, and the neuromorphic hardware that may one day implement it.
Why this list? Every agent framework bolts on a vector store and calls it "memory." These papers show what memory actually requires: admission control, consolidation loops, forgetting mechanisms, typed multi-store architectures, and retrieval strategies that go far beyond cosine similarity.
Maintained by @tfatykhov · Built alongside Nous, a cognitive AI agent with typed multi-store memory.
| If you want to... | Start with |
|---|---|
| Understand the landscape | Memory Survey (taxonomy) → Anatomy (critical analysis) |
| Build agent memory | xMemory (why RAG isn't enough) → A-MAC (admission control) → Mem0 (production system) |
| Add forgetting/consolidation | SleepGate → CraniMem |
| Evaluate your memory system | StructMemEval → Anatomy (why benchmarks are broken) |
| Train memory with RL | AgeMem (tool-based ops) → EMPO² (exploration) → MemFactory (framework) |
| Understand memory privacy risk | ADAM (extraction attack) → FSFM (forgetting as defense) |
| Unify memory ↔ skills ↔ rules | Experience Compression Spectrum → Externalization |
| Make memory writes safe & reversible | MemTX (transactional belief commit) → ChronoMem (versioning + rollback) |
| Ship something this week | Implementations & Reference Code (curated repos, mechanism-by-mechanism) |
| Explore the neuroscience | Neuromorphic section |
| Symbol | Meaning |
|---|---|
| ✅ | Venue confirmed — verified via official proceedings, OpenReview, or arXiv metadata |
| 📄 | Self-reported — metrics from authors' own evaluation; exercise caution |
| ⚠️ | Unverified — plausible claim but not independently confirmed |
| 🔬 | Editor's synthesis — cross-paper connection identified by the maintainer, not claimed by original authors |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem | Tang | Aug 2026 | The first serious attempt to say what a memory is, formally: memory is a basis, knowledge is its span, and answerability is a coverage problem — a query is answerable exactly when some single item in the span covers it. Optimal memory becomes the capacity-constrained maximizer of expected coverage, tracing a utility–capacity frontier that serves as a common yardstick for comparing memory systems. Also treats noise (coverage vs precision when the store holds false claims) and formalizes continual memory as a sequential MDP — memory is the state, writing is the action, query-time utility is the delayed reward. Instantiated concretely on Homer's Odyssey. ⚠️ Preprint; solo-author. |
| Memory in the Age of AI Agents: A Survey | 47 authors | Dec 2025 | Unified taxonomy across 3 lenses: Forms (token/parametric/latent), Functions (factual/experiential/working), Dynamics (formation/evolution/retrieval). Distinguishes agent memory from RAG and context engineering. Hugging Face Daily Paper #1 ⚠️ (unverifiable — HF archive page returns 404). |
| Anatomy of Agentic Memory | Jiang, Li, Wei et al. | Feb 2026 | Structured taxonomy of Memory-Augmented Generation (MAG) systems. Critically analyzes empirical fragility: benchmark saturation, metric misalignment, backbone-dependent variance, overlooked latency costs. Evaluates 5 systems (LOCOMO, A-Mem, MemoryOS, Nemori, MAGMA). |
| The Landscape of Agentic RL for LLMs | Guibin Zhang et al. (Oxford, Shanghai AI Lab, NUS, UIUC, UCL) | Sep 2025 | Synthesizes 500+ works. Reframes LLMs as autonomous agents using POMDPs. Covers planning, tool use, memory, reasoning, self-improvement. |
| From Static Templates to Dynamic Runtime Graphs | Yue et al. (IBM Research) | Mar 2026 | Agentic Computation Graphs (ACGs) framework — distinguishes workflow templates, realized graphs, and execution traces. Organizes ~40 papers by when structure is determined. |
| Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers | Pengfei Du | Mar 2026 | Write–manage–read loop formalization. 3D taxonomy: temporal scope × representational substrate × control policy. Five mechanism families: context-resident compression, retrieval-augmented stores, reflective self-improvement, hierarchical virtual context, policy-learned management. Proposes vision of a "foundation model for memory control." |
| Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering | Zhou, Chai, Chen, Guo, Shan, Song, Xu et al. | Apr 2026 | Unified review across four externalized components: memory stores, reusable skills, interaction protocols, runtime harness. Argues modern agent capability comes from reorganizing the runtime around the model, not from changing weights. Connects memory literature with skill-discovery and protocol-design literatures that rarely cite each other. |
| Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents | Zhang, Wang, Cui, Qiu, Li, Zhu, He | Apr 2026 | Positions memory/skills/rules on a single compression axis (5–20× episodic, 50–500× procedural, 1000×+ declarative). Citation analysis of 1,136 references finds <1% cross-community citation between memory and skills literatures. Identifies the "missing diagonal" — no current system supports adaptive cross-level compression. |
| From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms | Luo, Tian, Cao, Luo et al. | May 2026 | ACL 2026 Findings ✅ — Maps the field's evolution from passive storage stores toward experience-centric, self-evolving memory. Companion paper-list: Evolving-LLM-Agent-Memory-Survey. |
| Graph-based Agent Memory: Taxonomy, Techniques, and Applications | Yang, Zhou, Xiao, Dong et al. | Feb 2026 | First focused taxonomy of graph-based agent memory: node/edge designs, construction strategies (entity-centric vs event-centric vs hybrid), retrieval (one-hop vs multi-hop vs subgraph), and update/forget operators. Useful companion to A-MEM, HeLa-Mem, GAM. |
| Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory | Orogat, Mansour | May 2026 | Argues record-level correctness (rows, embeddings, edges) cannot satisfy long-term memory's needs, causing four recurring failure modes: unregulated growth, missing semantic revision, capacity-driven forgetting, read-only retrieval. Formalizes Governed Evolving Memory (GEM) — state-level operators (ingestion, revision, forgetting, retrieval) with six correctness conditions on the trajectory, not individual records. Prototype MemState on a property-graph backend. Proves no record-level system can satisfy the conditions regardless of storage model. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Metis: Memory Foundation Model | Zhang, Guo, Sun, Zhang, Hao, Lin et al. (17 authors) | Jul 2026 | Challenges this list's prevailing external-module framing. Argues agent memory should be native to the backbone, and formalizes native memory as (a) a persistent, dynamically evolving memory state inside the model and (b) native memory procedures that store and use information through model computation. Metis equips a foundation model with a native memory state accessed via memory attention, acquired through mid-training on large-scale memory-specific data. Online memory maintenance is gradient-free — a memory update requires only a forward pass, and all learned weights remain frozen at inference. Project and model checkpoints released. ⚠️ Preprint. |
| Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability | Zhou, Yu, Wei, Wu, Ouyang, Jiao et al. (11 authors) | Jul 2026 | The first systematic study of the memory design that is actually deployed by default — a directory tree of markdown files the agent itself reads, writes, and reorganizes with generic file tools — which prior research largely passed over. Formalizes it as three roles around one memory filesystem (management, search, execution agents), unifying declarative memory and skills in a single store. The default's two working assumptions do not hold up: what organization reliably buys is search economy (organized stores roughly halve retrieval cost where material is large), not better answers — no agent measured converts organization into accuracy, and organization erodes as the store grows for all but the strongest management agent. Changing the tool set alone reshapes the store as strongly as swapping the model 📄. ⚠️ Preprint. |
| MAGE: Memory as Agent-Guided Exploration | Chen, Lai, Feng, Han, Zhang, Lu, Li et al. | Jun 2026 | Argues semantic-similarity organization mismatches execution-state dependencies on long-horizon tasks — it fragments decision trajectories and mixes valid/erroneous traces. Stores interactions in a hierarchical state tree; agent state derives from the active root-to-current path (subgoal summaries + recent traces + branch hints). Four coupled ops: Grow, Compress, Maintain, Revise. +7.8–20.4pp task success on MemoryArena, -55.1% tokens 📄. |
| Mechanistic Attention Guidance for Agent Memory Refinement (AGMR) | Hong, Qiu, Wang, Yao | Jul 2026 | Existing self-evolving memory only inspects textual outputs (trajectories, reflections), risking unreliable error attribution and hallucinated edits. Uses retrieval-head attention as a mechanistic signal — aggregates attention over memory segments × decision steps into a context-utilization matrix that reveals actual memory-use patterns. AGMR corrects/enhances memory on failure, simplifies it on success, and re-executes to verify each update. Beats text-only refinement baselines on interactive decision-making benchmarks. Code released. ⚠️ Preprint. |
| Supra Cognitive Modes (SCM): A Routed Architecture for Agent Memory | Tobkin, Yang | Jul 2026 | Routes each query — factual lookup, relation-chain/current-state reasoning, or broad long-history synthesis — to a matched retrieval+synthesis payload over one shared ingest substrate (multi-granularity embeddings, extracted triples, fact-version metadata). A frozen semantic classifier + runtime gates dispatch among fused lexical/dense lookup, graph/multi-hop, and stratified long-form synthesis. Reports 84.87% LoCoMo factoid / 68.61% adversarial abstention, 61.49% MAB, 86.00% LongMemEval — but the authors themselves flag causal routing effects and efficiency gains as outside the available evidence 📄 (unusually candid self-limitation disclosure). ⚠️ Preprint. |
| A-MEM: Agentic Memory for LLM Agents | Xu, Liang, Mei, Gao, Tan, Zhang | Feb 2025 | Dynamic memory organization inspired by Zettelkasten. Creates interconnected "memory notes" with content, metadata, and explicit links. Agent autonomously decides when to create, update, or link memories. NeurIPS 2025 ✅. |
| SYNAPSE: Episodic-Semantic Memory via Spreading Activation | Jiang, Chen, Pan et al. | Jan 2026 | Memory as a dynamic graph where relevance emerges from spreading activation rather than pre-computed vector links. Lateral inhibition and temporal decay highlight relevant sub-graphs. Triple Hybrid retrieval. |
| MAGMA: Multi-Graph Agentic Memory Architecture | Jiang, Li, Li, Li | Jan 2026 | Multi-graph based architecture separating different memory concerns into distinct graph structures for AI agents. ACL 2026 Main ✅. |
| Mem0: Production-Ready AI Agents with Scalable Long-Term Memory | Chhikara, Khant, Aryan, Singh, Yadav | Apr 2025 | Scalable memory-centric architecture with dynamic extraction, consolidation, and retrieval. Graph-based variant for complex relational structures. 26% improvement over OpenAI memory, 91% lower latency, 90%+ token savings 📄 (self-reported). Evaluated on LOCOMO. |
| CraniMem: Neurocognitively Motivated Gated & Bounded Memory | Mody, Panchal, Kar, Bhowmick, Karani | Mar 2026 | Cranial-inspired dual-store: bounded FIFO episodic buffer + knowledge graph (long-term). Goal-conditioned gating. Utility tagging (Importance + Surprise + Emotion). Beats Mem0 by +57.6% on noisy HotpotQA 📄 (self-reported; "noisy" = authors' own distractor injection, not standard benchmark). ICLR 2026 MemAgents Workshop ✅. Code |
| AdaMem: Adaptive User-Centric Memory | — | Mar 2026 | Working, episodic, persona, and graph memories. Key innovation: question-conditioned retrieval planning — resolves target participant first, builds retrieval route combining semantic + relation-aware graph expansion. SOTA on LoCoMo and PERSONAMEM. |
| PlugMem: A Task-Agnostic Plugin Memory Module | Yang, Galley, Wang, Gao, Han, Zhai (Microsoft Research, UIUC) | Mar 2026 | Converts raw interactions into propositional (facts) + prescriptive (skills) knowledge, organized in a knowledge-centric memory graph. Knowledge — not entities or chunks — is the unit of memory access. Single plug-and-play module outperforms task-specific designs across 3 diverse benchmarks. Highest information density: more useful info per context token consumed 📄. |
| Memora: Harmonic Memory Representation | Xia, Zhang, Dixit, Harimurugan, Wang, Ruhle, Sim, Bansal, Rajmohan (Microsoft Research) | Feb 2026 | Proves that RAG and KG memory are special cases of this unified framework. Primary abstractions index concrete memory values; "cue anchors" expand retrieval beyond semantic similarity. New SOTA on both LoCoMo AND LongMemEval 📄. ICML 2026 ✅. |
| MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning | Lin, Zhang, Lu, Liu, Tang, He, Zhang, Wang (Microsoft Research) | Mar 2026 | Multi-agent framework: Meta-Thinker → Memory Manager → Query Reasoner. Backward path innovation: synthesizes probe QA pairs, verifies memory, converts failures into repairs BEFORE finalizing. Plug-and-play — improves 3 different storage backends on LoCoMo. Code |
| H-Mem: Hybrid Multi-Dimensional Memory | Ye, Huang, Chen, Zhang (Rutgers) | Mar 2026 | Organizes memory across time AND topic dimensions simultaneously. Mimics associative + hierarchical properties of human memory. EACL 2026 ✅. |
| H-MEM: Hierarchical Memory for High-Efficiency Long-Term Reasoning | Sun et al. | Mar 2026 | Multi-level memory storage with positional index encoding of sub-memory at each layer. Confidence-weighted retrieval — attaches memory weights to provide LLMs with uncertainty reference. Key finding: retrieval is ineffective without structured hierarchical storage. EACL 2026 ✅. |
| GAM: Hierarchical Graph-based Agentic Memory for LLM Agents | Wu, Zhang, Lin, Xu, Xu, Chen, Zou et al. | Apr 2026 | Explicitly decouples encoding from consolidation: an event-progression graph captures stream updates online; integration into the topic-associative network is deferred until a semantic shift is detected. Addresses the tension between stream-based fluidity and structured retention. |
| HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents | Zhu, Li, Zhang, Liu, Yang | Apr 2026 | ACL 2026 ✅ — Bio-inspired dual-graph: (1) episodic graph evolves via Hebbian co-activation; (2) semantic store populated via Hebbian Distillation — a Reflective Agent identifies densely-connected hubs and distills them into reusable semantic knowledge. Beats prior SOTA on LoCoMo across 4 categories with fewer context tokens 📄. Code |
| LightMem: Lightweight LLM Agent Memory with Small Language Models | Zhang, Zhang, Chen, Huang, Zheng et al. | Apr 2026 | ACL 2026 ✅ — SLM-driven memory with strict online/offline separation. STM/MTM/LTM tiers with two-stage retrieval (vector coarse → semantic re-rank). ~2.5 F1 over A-MEM on LoCoMo, 83ms retrieval, 581ms end-to-end 📄. Shows that careful SLM use can replace repeated large-model memory calls. |
| MemMachine: Ground-Truth-Preserving Memory for Personalized AI Agents | Wang, Yu, Love, Zhang, Wong, Scargall, Fan et al. | Apr 2026 | Open-source system integrating short-term, long-term episodic, and profile memory in a ground-truth-preserving pipeline. Targets multi-session degradation in standard RAG. |
| Omni-SimpleMem: Autoresearch-Guided Lifelong Multimodal Memory | Liu, Ling, Qiu, Liu, Han, Xia, Tu et al. | Apr 2026 | Uses autonomous research-agent search over the design space (architecture × retrieval × prompts × data pipeline) to discover effective lifelong multimodal memory configurations. The first paper to treat memory architecture itself as a search target. |
| Human-Inspired Memory Architecture for LLM Agents | Kerestecioglu, Robsky, Vasters, Sharma, Kesselman (Microsoft) | May 2026 | Biologically-grounded architecture with six cognitive mechanisms: sleep-phase consolidation, interference-based forgetting, engram maturation, reconsolidation on retrieval, entity knowledge graphs, hybrid multi-cue retrieval. Introduces a synthetic calibration methodology that derives all thresholds without benchmark exposure — eliminates a common eval-leakage source. First streaming M-tier LongMemEval eval (475 sessions). Dedup-based consolidation: 97.2% retention precision, 58% store reduction (+21.8 pp) on a 13K-issue VSCode dataset 📄. |
| MRAgent: Memory is Reconstructed, Not Retrieved — Graph Memory for LLM Agents | Ji, Li, Hooi (NUS) | Jun 2026 | ICML 2026 ✅ — Represents memory as a Cue-Tag-Content graph where associative tags bridge fine-grained cues to contents. An active reconstruction mechanism folds LLM reasoning into memory access — iteratively exploring and pruning retrieval paths on accumulated evidence, adapting retrieval to the reasoning context while avoiding combinatorial blowup. Up to 23% over strong baselines on LoCoMo + LongMemEval with substantially lower token/runtime cost 📄. The cleanest articulation of the "reconstruction ≠ retrieval" paradigm. |
| MemPro: Agentic Memory Systems as Evolvable Programs | Liu, Wang, Wu, Huang, Tao, Song, Zhou, He (ECNU) | May 2026 | Treats the entire memory construction-retrieval (MCR) pipeline as an evolvable program, not just the memory bank. Maintains a version tree of runnable memory-system implementations; an Evolving Agent selects promising versions, diagnoses recurring failures, and synthesizes improved children via failure-mode-guided edit-debug refinement. Beats static and prompt-level evolving baselines on LongMemEval / LoCoMo / HotpotQA / NarrativeQA within a few iterations, with favorable performance-cost trade-off 📄. Code available. ⚠️ Preprint; venue unconfirmed. |
| MemDreamer: Hierarchical Graph Memory with Agentic Retrieval for Long Video | — | Jun 2026 | Extends agent memory to the multimodal long-video setting: a hierarchical graph memory plus agentic retrieval over hours-long video, exploiting the tree+graph structure for question answering. ⚠️ Preprint; venue unconfirmed. |
| H-Mem (Yu & Fang): Evolving & Retrieving Agent Memory via a Hybrid Structure | Yu, Fang, Liu, Ma (CUHK-Shenzhen / Huawei) | May 2026 | Distinct from the EACL H-Mem (Rutgers) and the Sun et al. H-MEM listed above — a third system sharing the name. Hybrid tree + graph: a temporal-semantic tree lets short-term memory evolve into a summarizing long-term store, while a parallel knowledge graph captures entity relationships; retrieval exploits both. SOTA on QA across three agent-memory benchmarks 📄. ⚠️ Preprint; venue unconfirmed. |
| MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management | Liu, Wu, Liu, Zhao, Liu, Li, Zhang, Wang, Guo, Liu | Jun 2026 | Extends agent memory to long-horizon mobile GUI agents. Introduces Context-as-Action (ConAct): context management becomes a first-class action emitted by the same policy that selects UI actions, maintaining three structured fields (folded action history, folded UI state, recent step record) instead of passively appending ReAct-style transcripts. 8B model trained on the 2,956-trajectory MemGUI-3K dataset achieves best open-data 8B performance on MemGUI-Bench and generalizes OOD to MobileWorld. Code/data to be released. ⚠️ Preprint. |
| Multi-Head Recurrent Memory Agents | Li, Yeh, Li | Jul 2026 | Diagnoses why recurrent memory agents degrade at long context: decomposing performance into capture vs retention shows retention collapse is the dominant bottleneck, caused by treating memory as one monolithic text block that every update risks overwriting. Multi-Head Recurrent Memory (MHM) partitions memory into independent heads with a stage-wise select-then-update strategy — only one head updates per step, the rest are structurally shielded. The MHM-LRU instantiation is training-free and lifts RULER-HQA retention at 896K tokens from <30% to 73.96%. Architecture, not model behavior, is the lever. 📄 ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| When Not to Write Memory: Governing False Promotion from Correlated Agent Traces (GovMem) | Qi, Xu, Li | Jun 2026 | Reframes admission as a write-path governance problem: repeated observations aren't independent evidence if copied from a shared source, induced by a shared prompt, or valid only in a narrower scope. GovMem estimates dependency-aware support, retrieves counterevidence, assigns scope, and outputs promote / reject / needs-review. Cuts false promotion 0.597→0.040 (synthetic) at 0.960 recall; on a human-labeled real-trace subset, false promotion falls to 0.032 but held-out false promotion stays 0.111 at 0.692 review burden. On 133 high-impact external coding-agent candidates, none were judged safe for automatic promotion. Positioned as a diagnostic governance design point, not a validated auto-writer — a sobering result for anyone auto-promoting agent traces into memory 📄. Accepted at MLISE 2026 ✅. |
| A-MAC: Adaptive Memory Admission Control | Zhang et al. (Workday AI) | Sep 2025 | 5-dimension scoring: Utility (LLM call), Confidence (ROUGE-L grounding), Novelty (1 - max cosine similarity), Recency, Type Prior. Learned weights vs fixed heuristics. |
| ACC: Agent Cognitive Compressor | Bousetouane | Jan 2026 | Bio-inspired memory controller replacing transcript replay with bounded internal state updated online. Compresses agent cognitive state without losing decision-relevant context. |
| MemReader: From Passive to Active Extraction for Long-Term Agent Memory | Kang, Li, Chen, Tang, Xiong, Li | Apr 2026 | First RL-trained active extraction policy (vs passive transcription). MemReader-4B uses GRPO + ReAct to evaluate value/ambiguity/completeness, then chooses WRITE / DEFER / RETRIEVE-CONTEXT / DISCARD before admission. Integrated into MemOS. Addresses memory pollution from noisy dialogue and cross-turn dependencies. |
| Personalize-then-Store: Benchmarking & Learning Personalized Memory for Long-horizon Agents | In, Kim, Park, Yoon, Park (KAIST) | May 2026 | Argues universal static storage policies waste budget on transient sessions while dropping critical long-horizon context. Introduces PerMemBench (first benchmark for personalized memory policies, with multi-year multi-domain personas) and session-level storage gating — a lightweight per-user admission filter that bypasses memory ops for transient sessions. Personalization yields large retention gains under perfect gating; accurate gating remains the open challenge 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory (ReFind) | Li, Zhang, Xu, Du, Fu, Chen | Aug 2026 | Asks how much of structured memory's benefit comes from the structure versus from competent retrieval over the raw history — and answers uncomfortably. ReFind builds no semantic structure at all: it leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls (session-aware rank fusion, local context expansion, temporal narrowing, skipping already-inspected sessions). Across ~2,800 questions under MemoryAgentBench's incremental multi-turn setting it attains the highest mean accuracy (58.2) of any system compared — above the strongest graph- and tree-based systems (HippoRAG 2, 53.2) — all under a GPT-4o-mini backbone matched to every reused baseline. Reaches 93.2±3.3 / 89.3±6.0 on LongMemEval-S/M with GPT-5-mini, with no LLM-based index construction at all 📄. ⚠️ Preprint. |
| xMemory: Beyond RAG for Agent Memory | Hu, Zhu, Yan, He, Gui | Feb 2026 | 4-level memory hierarchy. Key thesis: RAG ≠ agent memory. Submodular diversity-aware retrieval (MMR) replacing naive top-k. Uncertainty-gated adaptive expansion. Theme clustering. 44.91% retroactive reassignment rate proves flat structures fail. ⚠️ Conference status unverified. |
| SuperLocalMemory V3 | — | Mar 2026 | First information-geometric foundations for agent memory. Fisher information metric replaces cosine similarity. Riemannian Langevin dynamics for retrieval. Theoretical grounding for memory operations. |
| ExpRAG: Retrieval-Augmented LLM Agents Learning to Learn from Experience | Ferraz, Deffayet, Nikoulina, Déjean, Clinchant | Mar 2026 | Trains agents to USE retrieved trajectories in-context (retrieval-augmented fine-tuning). Standard LoRA collapses on OOD tasks; ExpRAG-LoRA generalizes to held-out hard tasks. Combines experience retrieval with fine-tuning — neither alone is sufficient. |
| To Know is to Construct: Schema-Constrained Generation for Agent Memory (SCG-MEM) | Zheng, Song, Li, Yang | Apr 2026 | Replaces dense retrieval with schema-constrained decoding. Piaget-inspired: assimilation (grounding into existing schemas) vs accommodation (expanding schemas with novel concepts). Solves "structural hallucination" — LLMs generating references to non-existent memory keys. Constructivist alternative to similarity-based recall. |
| SAM: State-Adaptive Memory for Long-Horizon Agents | — | May 2026 | Targets retrieval of scattered information over long horizons — adapting memory access to evolving task state rather than fixed top-k recall. ⚠️ Preprint; venue unconfirmed. |
| RaMem: Contextual Reinstatement for Long-term Agentic Memory | Yang, Kan, Li, Li, Qin, Li, Bogdan, Thomason | Jun 2026 | Names and addresses context collapse: retrieved memory fragments that share entities or user states can look equally relevant even when their surrounding episodic conditions (time, session, participants) differ, so similarity-based retrieval returns content-relevant but context-invalid evidence. RaMem runs four coordinated stages — evidence anchoring, recall-condition induction, validity-aware retrieval, context-preserved synthesis — to prioritize context-compatible memories over merely similar ones. +10% average F1 over strong baselines across several backbones on long-term memory benchmarks 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| TEPA: Revoking Stale Memories for Conflict-Robust Language Agents | Zhou, Ouyang, Zheng, Xiang | Aug 2026 | Makes validity an explicit state of memory. Observations are stored as keyed precedents; an active precedent is revoked when fresh evidence contradicts it under the same key, so retrieval draws only from current evidence while revoked history is preserved for audit rather than deleted. The headline result is a failure mode for the naive strategies: under full reversal, append-only and last-write-wins both score 0.210 — below the 0.309 of having no memory at all — while TEPA reaches 0.950 (50 seeds; reproduced under real file-backed execution at 0.203 / 0.298 / 0.950). Honest scoping: on clean MemoryAgentBench SH-6k, TEPA only matches a strong last-write-wins cache, confirming current-key replacement as the decisive op for single-hop facts; multi-hop and very-long-context settings expose retrieval-chain bottlenecks beyond fact-level validity tracking 📄. ⚠️ Preprint. |
| Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory | Kang, Liu, Kai, Liang, Tang et al. | Jul 2026 | Reframes retention-vs-consolidation as a budget-dependent operator choice, not an architectural commitment. Decomposes each operator's utility (Merge / Abstract / Rewrite) into a coverage effect on evidence that retention omits plus a signed replacement effect on raw evidence that already fits — their balance explains why the preferred action changes with relative budget pressure. Implemented as Offline Abstraction-Safety (OAS), a lightweight learner estimating action utility from pre-generation features with held-out harm calibration. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets while retention is preferable under loose ones; LoCoMo replicates the crossover at a smaller budget, consistent with its shorter evidence. Cross-note abstraction and merging generally outperform local rewriting 📄. ⚠️ Preprint. |
| Control-Plane Placement Shapes Forgetting | Yang | Jun 2026 | 13-configuration architectural study of where an LLM should sit in the memory pipeline — recall plane (extensively benchmarked) vs. control plane that mutates via supersede/release/purge (largely untested). Finding: deterministic primitives handle lexical/temporal forgetting but fail canonicalization (5% identifier-obfuscation, 0% cross-lingual); inscribe-time LLM fixes canonicalization (100%) but can't do intent-aware deletion (0%); a mutation-time hook recovers intent-aware deletion (78–85%) and lifts nearly every category simultaneously (91.7–93.2% overall) at $0.17/385-case run. Releases ForgetEval (1000-case + 385-case adversarial suite, 10-annotator Fleiss' κ=0.958) plus a 6-method Adapter Protocol. Core claim: production failures are predominantly forgetting failures, not recall failures — yet benchmarks measure only recall. Code/benchmark released (MIT) 📄. |
| SleepGate: Sleep-Inspired Forgetting for LLMs | Xie (Kennesaw State) | Mar 2026 | Learned sleep cycle over KV cache addressing proactive interference. Conflict-aware temporal tagger + forgetting gate + consolidation module. Reduces interference horizon from O(n) to O(log n). 99.5% retrieval accuracy at PI depth 5 vs 23% baseline 📄 (self-reported; extraordinary gap warrants independent replication). Supersession detection via binary σ flag. |
| SCM: Sleep-Consolidated Memory with Algorithmic Forgetting | Shinde | Apr 2026 | Five components inspired by human memory: limited-capacity working memory, multi-dimensional importance tagging, offline sleep-stage consolidation with distinct NREM and REM phases, intentional value-based forgetting, and a computational self-model for introspection. Reports perfect 10-turn recall, 90.9% noise reduction, sub-millisecond search 📄. Research preview. |
| FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory | Gu, Xiong, Wang, Ren, Li, Zhang, Guo et al. | Apr 2026 | Grounds selective forgetting in hippocampal indexing/consolidation theory and the Ebbinghaus curve. Argues that in resource-constrained environments a forgetting mechanism is as crucial as retention — and that memory security depends on the agent's ability to actively drop sensitive history. |
| Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning (LRE) | Lia, Mazumder | Jun 2026 | Reframes eviction as a fidelity problem, not a compression problem: dropping one load-bearing detail (an access token, a required path) fails the task outright. Learned Relevance Eviction (LRE) is a few-kilobyte, CPU-only, LLM-free scorer that learns which history units are load-bearing and keeps them verbatim. Matches full-history-retention accuracy while cutting peak context up to 52%; on LoCoMo gives the best budgeted answer quality reading 68% fewer tokens; annotation-free training on the system's own behavior recovers 95% of supervised effectiveness. Evidence that cheap learned relevance can replace LLM-mediated summarization for eviction 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| AgeMem: Agentic Memory — Learning Unified LTM and STM Management | Yu, Yao, Xie, Tan, Feng, Li, Wu | Jan 2026 | Exposes store/retrieve/update/summarize/discard as tool-based actions. Agent learns WHEN and WHAT to do via 3-stage progressive RL + step-wise GRPO. ACL 2026 SAC Highlight ✅ — Outperforms all heuristic-based baselines across 5 benchmarks. Key thesis: memory operations should be learned, not hard-coded 📄. |
| EMPO²: Exploratory Memory-Augmented On- and Off-Policy Optimization | Liu, Kim, Luo, Li, Yang | Feb 2026 | Uses memory not just for recall but for exploring novel states. Hybrid on/off-policy RL — 128.6% improvement over GRPO on ScienceWorld 📄. Adapts to new tasks with just a few memory-augmented trials, zero parameter updates. ICLR 2026 ✅. |
| MemFactory: Unified Inference & Training Framework for Agent Memory | Guo, Li, Tang, Xiong, Li | Mar 2026 | "LLaMA-Factory for memory agents" — first unified modular framework. Lego-like plug-and-play memory components. Natively integrates GRPO for RL-based memory policy training. Supports Memory-R1, RMM, MemAgent paradigms out of the box. Up to 14.8% improvement over base models 📄. |
| Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents | Yan, Bahloul, Nie, Schwarzmann, Trivisonno, Tresp, Ma | May 2026 | Identifies a fundamental flaw in GRPO for memory-RL: once rollouts write different memories, they no longer share the same effective environment, so trajectory-level group comparisons are unfair. Introduces LoGo-GRPO (Local + Global) — global keeps end-to-end long-horizon reward; local re-rollouts compare memory-op outcomes from the same intermediate state. Shared-parameter co-learning for fact extractor + memory manager. Progressive curriculum 8→16→32 sessions. |
| What Training Data Teaches RL Memory Agents: Curriculum Effects in Memory-Augmented QA | He, Lin, Liu, Wu, Xie, Zhou, Xiao | May 2026 | Controlled study holding architecture/RL/hyperparams fixed and varying only curriculum across in-domain (LoCoMo), mixed (LoCoMo+LongMemEval), and OOD (LongMemEval). Key findings: curriculum is a fine-grained lever on specialization, not a uniform scaling factor; per-type differences dwarf aggregate differences (single-number benchmark comparisons systematically underreport); binary EM reward produces no signal at G=4 group size — continuous rewards needed for single-GPU regime. Practical RL-memory training playbook. Code |
| SaliMory: Orchestrating Cognitive Memory for Conversational Agents | Zhang, Zhang, Jiang, …, X. L. Dong (Meta) | Jun 2026 | Trains a single LM to manage a cognitively-structured memory (user facts, preferences, working memory). Introduces a hierarchical stage-wise process reward + reward-decomposed contrastive (GRPO-style) refinement, giving isolated supervision to distinct ops (selective filtering, consolidation, cue-driven recall) end-to-end — sidestepping the credit-assignment bottleneck of standard RL over multi-stage pipelines. Cuts memory-attributed failures by ~⅓, +10% end-to-end accuracy, more than doubles the Good-Personalization rate 📄. ⚠️ Preprint. |
| JAMEL: Joint Agent Memory & Exploration Learning via Novelty Signals | Tian, Weng, Kong, …, Y. Li (Tsinghua / AIR) | Jun 2026 | Co-trains the memory module and exploration policy in a mutually-dependent loop: exploration needs memory to tell exhausted from unseen behaviors, while novelty-seeking interaction supplies annotation-free supervision (e.g., code coverage) for the latent memory. Generalizes to unseen environments; rivals a closed-source model while reducing tokens 📄. Code/model open-sourced. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| The Missing Memory Hierarchy: Demand Paging for LLM Context Windows | Mason (UBC / Georgia Tech) | Sep 2025 | Maps OS virtual memory concepts to LLM context: physical memory = context window, virtual memory = persistent state, page table = retrieval handles, page fault = re-request evicted content. Analyzed 857 sessions, 54,170 API calls, 4.45B effective input tokens. |
| What to Keep, What to Forget: A Rate–Distortion View of Memory Compaction in LLMs and Agents | Colaco, Lahjouji | Jul 2026 | Unifies four separate compaction literatures — KV-cache eviction/quantization, prompt pruning/distillation, bounded architectural state, and agent memory consolidation — as instances of one rate–distortion problem: what to retain vs discard, at what fidelity, under a resource budget, to preserve downstream utility. Builds a layer-agnostic lower bound and a seven-axis taxonomy that lets mechanisms transfer between layers that have never been connected (serving-stack KV management ↔ agent long-term memory). Finding: at every layer the keep/discard signal is attention magnitude or recency, and it fails the same way everywhere — discarding before the query is known, with no way to undo it. Proposes a cross-layer benchmark that no existing benchmark provides. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents | Huang, Zhang, Wu, Chen, Jiang, Yang et al. — incl. Ying Nian Wu, Kai-Wei Chang, Philip S. Yu, Aylin Caliskan (15 authors) | Aug 2026 | Controlled harness comparing memory substrates — the underlying medium memory is represented in: dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms — across 3 backbones × 4 benchmark suites under 26 performance and efficiency metrics. Headline: no single substrate consistently dominates. Broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context; substrates that do well at moderate history lengths become costly or brittle at longer horizons. Motivates substrate routing as a necessary component of adaptive agent memory. ⚠️ Preprint; code on acceptance. |
| Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory | Li, Du, Xu, Mao | Jul 2026 | Names an assumption so natural it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks it — a tree-nut allergy should change the answer to a macaron request through almond flour, yet the two texts share no cue a retriever can see. InMind is a 125-task, expert-verified benchmark across ten life domains (113 tasks grounded in citable public sources) whose paired controls separate three explanations existing evaluations conflate: never stored, no bridging knowledge, or stored-but-never-surfaced. The verdict is clean: with the decisive memory placed in context the backbone answers 84.0% of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic systems reach at most 14.4% — even though they recall those same facts on demand at up to 100%. An embedding with 8× the dimensionality does not close the gap. Locates the failure in the query-conditioned interface itself, naming routing — deciding which facts stay visible — as the open problem. |
| MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation | Wang | Jun 2026 | Controlled one-variable-at-a-time protocol on LongMemEval-S across 3 model families, exposing confounds that inflate architecture claims: verbatim RAG ≈ full-context GPT-4o-mini (47.2% vs 49.8%, p=0.34) but the ranking reverses by model (Gemini +14pp from full context; Sonnet +31pp from RAG, partly by refusing 63% of full-context queries); swapping only the embedding model shifts accuracy +6.2pp and flips which system wins; agent self-memory (42%) underperforms basic retrieval (47%); Mem0 matches cloud RAG on only 2/6 question types at 50x the cost. Recommends fixing embedding models across comparisons and reporting write-path cost before attributing gains to architecture — directly relevant to how Infini Memory/MAB-style comparisons should be read 📄. |
| StructMemEval: Evaluating Memory Structure in LLM Agents | Shutova, Olenina, Vinogradov, Sinitsin | Feb 2026 | Tests agents' ability to organize memory (not just retrieve). Tasks: transaction ledgers, to-do lists, trees. Key finding: LLMs don't spontaneously recognize when to apply memory structure — they succeed only when explicitly prompted. |
| MemoryAgentBench: Evaluating Memory via Incremental Multi-Turn Interactions | Hu, Wang, McAuley (UC San Diego) | Jul 2025 (revised Jun 2026) | Tests 4 memory competencies in realistic incremental accumulation. Key finding: no current method masters all 4 competencies simultaneously. ICLR 2026 ✅. Code |
| From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (Memora + FAMA) | Uddin, Shubham, Blanco, Baral, Wang (ASU) | Apr 2026 | ACL 2026 Findings ✅ — Introduces Memora, a weeks-to-months benchmark over three tasks: remembering, reasoning, recommending. Introduces FAMA (Forgetting-Aware Memory Accuracy) — a metric that penalizes reliance on obsolete or invalidated memory rather than just rewarding recall. Evaluation of 4 LLMs and 6 memory agents finds frequent reuse of invalid memories and failures to reconcile evolving knowledge. ⚠️ Not to be confused with Microsoft's Memora: Harmonic Memory Representation (2602.03315). |
| STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? | Chao, Bai, Sheng, Li, Sun | May 2026 | Benchmark for Implicit Conflict — a later observation invalidates an earlier memory without explicit negation, requiring inference + commonsense to detect. 400 expert-validated scenarios, 1,200 queries, contexts up to 150K tokens. Three-dimensional probing: State Resolution, Premise Resistance, Implicit Policy Adaptation. Best frontier LLM only 55.2% overall. Companion prototype CUPMem uses structured state consolidation + propagation-aware search. Complements FAMA. |
| MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks | He, Wang, Zhi, Hu, Chen, Yin, Wu, Ouyang, Wang, Pei, McAuley, Choi, Pentland | Feb 2026 | Memory-Agent-Environment loop benchmark across web navigation, preference-constrained planning, progressive information search, sequential formal reasoning. Key result: agents near-saturated on LoCoMo perform poorly in MemoryArena, exposing a gap between long-context memorization benchmarks and interdependent agentic memory use. Project |
| MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery | Ma, Zhou, Huang, Yang, Ma, Wang, Li, Miao, Yu, Wang | Jun 2026 | Argues memory should be evaluated as an auditable post-interaction artifact, not just through downstream task success. Simulates 50 users each carrying a hidden, taxonomy-anchored 31-dimension state bank across leak-controlled tasks, then reconstructs that state from the agent's resulting memory (full-store and top-k access) and scores it against ground truth. Key finding: task completion nearly saturates even for a memoryless baseline, while category-balanced state recovery stays moderate (~0.6) and drops further under top-k — successful assistance and recoverable memory are distinct capabilities. First benchmark to score memory recovery directly. |
| Are We Ready For An Agent-Native Memory System? | Zhou, Zhou, Han, Xu, Li, Li, Xiong, Wu | Jun 2026 | Systematic data-management-perspective evaluation decomposing agent memory into four modules — representation/storage, extraction, retrieval/routing, maintenance — and benchmarking 12 representative memory systems across 5 workloads / 11 datasets. Key finding: no single architecture dominates; effectiveness depends on whether the memory structure matches the workload bottleneck. Fine-grained ablations quantify effects on representation fidelity, retrieval precision, update correctness, long-horizon stability; localized maintenance beats global reorganization on cost. Code and paper-list released. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems | Bhandari, Wadhwani, Kumar, Narang | Aug 2026 | Takes one specific mechanism from western scrub jay episodic memory — per-memory, type-conditioned temporal decay — and operationalizes it as an auto-classified perishability coefficient in an external store, against the default that all memories are equally persistent. Each memory is a jointly-bound What–Where–When tuple carrying estimated perishability and a utility horizon, retrieved by query-adaptive scoring and revised retroactively at O(1) LLM calls per update. Introduces the Temporal Generalization Test (held-out retention intervals) and a Generalization Gap metric, on which ScrubJay-MEM is the only retrieval-based system with substantially positive GenGap (+0.108); on MemoryAgentBench EventQA-64k it improves F1 by +2.66 over Mem0 and +3.09 over Qwen3-Embedding-4B, and a decay ablation collapses GenGap by 5.7×. Unusually candid scoping: gains narrow under stronger backbones and reverse on fact-consolidation tasks 📄. ⚠️ Preprint. |
| Episodic Memory is the Missing Piece for Long-Term LLM Agents | Pink et al. (UT Austin) | Feb 2025 | Position paper arguing episodic memory — supporting single-shot learning of instance-specific contexts — is the critical missing capability. Defines 5 properties: temporal, instance-specific, single-shot, inspectable, compositional. |
| MAP: Modular Agentic Planner | Correa, Samwick, Gershman et al. (Microsoft Research, Harvard) | Oct 2023 / Nature Comms 2025 | Brain-inspired architecture decomposing planning into PFC-associated modules: conflict monitoring, state prediction, state evaluation, task decomposition, task coordination. arXiv:2310.00194. |
| SCL: Structured Cognitive Loop with Governance Layer | Kim | Nov 2025 | R-CCAM: Retrieval → Cognition → Control → Action → Memory. Soft Symbolic Control = governance layer applying symbolic constraints to probabilistic inference. Zero policy violations, complete decision traceability. |
| Procedural Memory Is Not All You Need | Wheeler, Jeunen | May 2025 | LLMs are constrained by reliance on procedural memory (pattern-driven tasks). For "wicked" learning environments with shifting rules and ambiguous feedback, different memory types are required. ACM UMAP '25. |
| Evaluating Theory of Mind in Multi-Agent LLM Systems | Kostka, Chudziak (Warsaw UT) | Sep 2025 | ToM and Internal Belief mechanisms are NOT universally beneficial — stronger models handle extra cognitive load well, but weaker models get confused. Model capability is the dominant factor. ICCCI 2025 ✅ (journal_ref). |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Demystifying Agent Skills: Why They Work—Until They Don't | Jiang, Huang, Xing, Wu, Gao et al. | Aug 2026 | Moves skill evaluation past aggregate success rates by isolating representation, outcome annotation, retrieval difficulty, and cross-framework robustness. Normalizes 8,135 trial records into a taxonomy of three categories and twelve skill-use modes. Core mechanism finding: skills work because noisy trajectories become procedural anchors that stabilize execution — 65.7% of cases, versus 4.5% for explicit knowledge injection (+6.06 points over Workflow Memory in matched comparisons); skills stabilize action, they do not inject missing facts. The scaling warning is the more useful half: retrieval is a separate bottleneck — as pools grow from 5 to 100, actual-use precision falls 29.6% → 3.3%. Exact ground-truth invocation is found to be neither sufficient nor necessary. Directly relevant to anyone scaling a skill/procedure catalog 📄. ⚠️ Preprint. |
| Muscle Memory for Agents: Compile not Merely Retrieve | Omran, Lanka, Zhang, Dixit | Aug 2026 | Position paper against the field's default pattern (store experience → retrieve at inference → let a general-purpose orchestrator interpret it), arguing it is the wrong default for personalization. Proposes compiling recurring user intent into purpose-built specialist agents as a memory paradigm distinct from retrieval, targeting the multi-turn tax where users repeatedly correct format, depth, and scope. Reference implementation is a four-phase pipeline (Harvest → Analyze → Augment → Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable specialists with two-stage trigger matching. On 90 held-out scenarios across five personas it wins 32 of 36 cases where a specialist fires (88.9%), at +2.05 personalization gain and a −0.28 accuracy cost on a 1–4 scale — note the win rate is conditional on a specialist firing 📄. ⚠️ Preprint. |
| SkillRouter: Skill-Based Routing for LLM Agents | Alibaba | Sep 2025 | Critical finding: skill BODY is the decisive routing signal (91.7% attention weight), NOT name (7.3%) or description (1.0%). Removing body causes 29-44pp degradation. BM25 on metadata alone scores 0%. Two-stage retrieve-and-rerank (1.2B params). |
| EvoSkill: Self-Evolving Skill Discovery | Sentient, Virginia Tech | Sep 2025 | Automates skill discovery via iterative failure analysis without model fine-tuning. Three agents: Executor, Proposer, Skill-Builder. Skill-merge outperforms single runs. Code |
| Trajectory-Informed Memory Generation for Self-Improving Agent Systems | Fang, Isahagian, Jayaram, Kumar, Muthusamy, Oum, Thomas (IBM Research) | Mar 2026 | Extracts typed actionable tips from execution trajectories: strategy tips (clean successes), recovery tips (failure-then-recovery), optimization tips (inefficient-but-successful). Closes the gap between episodic logs and reusable procedural knowledge — useful as a complement to EvoSkill and SkillRouter. |
| DeltaMem: Incremental Experience Memory for LLM Agents via Residual Trees | Tan, Zhang, Cao, Li, Chen (RUC) | Jun 2026 | Stores experience as residual deltas across two trees — goal-conditioned task skills and scene-level environment knowledge — so related episodes share a common root instead of duplicating content. Retrieval: failure-penalized similarity scan + root-to-match chain reconstruction; autonomous consolidation distills high-frequency paths into new roots. Cuts redundancy and retrieval conflicts; beats baselines across interactive environments 📄. Code released. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems | Xiong, Zhou, Wang, Gu, Tang, Li et al. (9 authors) | Aug 2026 | Targets a path defenses structurally miss: a poisoned memory is written once, retrieved repeatedly, promoted into shared memory, and reused by other agents — so a single write steers many later decisions and contaminates agents that never saw the original attack, while no malicious message ever crosses a visible communication edge at the moment of harm. Because existing safeguards inspect prompts, actions, or communication edges, they miss content that looks benign at write time but becomes harmful after retrieval. MAPLE-Guard monitors the memory lifecycle and places gates at write, retrieval, promotion, and cross-agent reuse, quarantining risky memories and blocking poisoned private memories before they enter shared memory. ASR 38.2% → 0.9% (LongMemEval) and 34.7% → 0.2% (AppWorld); multi-agent defense success rate 54.0%→74.3% and 42.5%→99.8% 📄. Code ⚠️ Preprint. |
| Governed Shared Memory for Multi-Agent LLM Systems (MemClaw / ArgusFleet) | Margalit, Cohen-Inger, Avram, Taig, Margalit | Jun 2026 | Formalizes the fleet-memory problem: unauthorized leakage, stale propagation, contradiction persistence, provenance collapse. Defines systems-level primitives — scoped retrieval, temporal supersession, provenance tracking, policy-governed propagation — implemented in MemClaw, a production multi-tenant memory service, evaluated live via ArgusFleet rather than a baseline comparison. Results: 100% correct depth-4 provenance reconstruction at sub-second/hop; zero cross-fleet leakage. Also reports real production bugs found via live eval: sub-tenant scope bypass on direct GET-by-id (disclosed/remediated) and a pipeline-ordering conflict where a synchronous near-duplicate gate can reject contradictory writes before the async contradiction detector runs. Conclusion: long-context retrieval alone is insufficient for production multi-agent memory 📄. |
| CoMAM: Collaborative Memory for Multi-Agent Systems | — | Sep 2025 | Collaborative RL framework modeling memory agents as sequential MDP with inter-agent dependencies. Group-level ranking consistency for coordinated memory operations. |
| DCM-Agent: Dual-Cluster Memory for Multi-Paradigm Ambiguity | Zhang, Wan, Zhang, Yang, Zhang, Wei, Liu | Apr 2026 | Tackles structural ambiguity in optimization problems where one problem admits multiple conflicting modeling paradigms. Training-free dual-cluster memory keeps competing paradigm solutions separated so the agent can pick or blend at inference time. Relevant to multi-agent settings where agents disagree on framing. |
April 2026 marks memory security emerging as a first-class research concern. As agent memories grow rich with user data, they become attack surfaces.
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| SkillJack: Persistent Skill Backdoors in Self-Evolving Agents | Ying, Wu, Wu, Zheng, Cheng et al. | Aug 2026 | Prior poisoning work only bites when a poisoned record is retrieved; SkillJack attacks the experience-to-skill pipeline instead, hijacking the agent's own learning process so poisoned experiences become durable behavioral artifacts. Three named properties: sanitization whitewashing (malicious intent obscured during skill extraction), cross-layer promotion (transient experience becomes persistent capability), and persistence isolation — 80.0% of skill-mediated attacks survive deletion of the original poisoned records, so source-record cleanup is not a remedy. On SkillX, safety detection drops from 98.5% on poisoned trajectories to 11.4% on the extracted skills, with attack success 56.2% (SkillX) / 89.2% (Anything2Skill) over 150 shared trajectories; some implanted skills unintentionally activate on benign queries. Motivates provenance-aware skill lifecycle protection. Code released (Tencent/AI-Infra-Guard) 📄. ⚠️ Preprint. |
| Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall (PPMF) | Xu, Xiao, Shao, Liu, Li | Jul 2026 | The attack-side twin of authority collapse. Identifies memory provenance laundering: during LLM-based consolidation, an external untrusted observation is rewritten as apparent user history or workflow support — preserving the action trigger while erasing the low-trust source that should limit its authority. Prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy consolidation. PPMF is lightweight middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. Vulnerable consolidated memories reach up to 1.000 ASR; with provenance, confirmation, and risk labels intact, no evaluated unauthorized high-risk action passes the gate while benign actions remain executable 📄. ⚠️ Preprint — arXiv comment reads "EMNLP2026 submitted" (submitted, not accepted). |
| MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair | Chen, Xie, Fu, Zhou, Yu, Xuan | Jul 2026 | Traces the same malicious semantics across persistence, downstream consequence, and selective repair — the repair axis prior poisoning benchmarks omit. 310 cases drawn from 48 realistic contexts (code/science, daily life, office work), each following a controlled Write–Execute–Forget protocol in an isolated runtime, with evidence-based adjudication across seven lifecycle checkpoints over a 24-configuration matrix (2 agent harnesses × 4 memory backends × 3 LLM backends). Across all 24: malicious memory persists in 84.2% of cases and the full Write–Execute chain succeeds in 50.3%; among successfully poisoned cases 59.6% complete the full Execute chain and 56.1% achieve selective repair 📄. ⚠️ Preprint. |
| Securing LLM-Agent Long-Term Memory Against Poisoning (TMA-NM) | Louck | Jun 2026 | Proves content- and lineage-based memory-poisoning defenses are malleable — an attacker can launder untrusted origin through the agent's own summarization, a trusted-tool echo, or manufactured corroboration, making poisoned content look benign and flipping its derivation edge to "trusted." Formalizes the malleability problem for the write-retrieve-act pipeline and proves a machine-checked separation theorem: write-time origin binding is necessary, non-malleable origin-bound authority with Sybil-resistant corroboration-gated elevation is sufficient. TMA-NM hits 0% attack success (vs. up to 68% laundering ASR for existing defenses) across 8 frontier models at full legitimate utility. Releases benchmark, harness, and machine-checked TLA+ models 📄. |
| When Agents Remember Too Much: Memory Poisoning Attacks on LLM Agents (GhostWriter / AM-Sentry) | Torres, Shrestha, Misra | Jul 2026 | Targets personal assistant agents — the convergence of conversational + action-planning agents handling sensitive data via untrusted sources. GhostWriter attacks in two phases (injection → activation) and achieves ~98% injection rate, ~60% activation rate against SOTA agents, exploiting the lack of security-focused memory governance. Proposes AM-Sentry (memory-saving policy + memory-retrieval screen) which dramatically cuts GhostWriter's success while preserving utility. |
| ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying | Lyu, He, Wang, Hu, Li, Chen, Li, Chen | Apr 2026 | First systematic privacy attack on agent memory. Combines data-distribution estimation of a victim agent's memory with entropy-guided adaptive querying to maximize leakage. Achieves up to 100% Attack Success Rate on extracting stored memories 📄 — substantially outperforming prior attacks. Establishes the need for privacy-preserving memory designs as a first-class research direction. |
| SSGM: Stability and Safety Governed Memory for LLM Agents | Lam, Li, Zhang, Zhao | Mar 2026 (rev May 2026) | Conceptual governance framework that decouples memory evolution from execution by enforcing consistency verification, temporal decay modeling, and dynamic access control before any memory consolidation. Provides a taxonomy of memory corruption risks: topology-induced knowledge leakage, semantic drift via iterative summarization, and consolidation hazards. Complementary defensive framing to ADAM. |
Wave 7 (June–July 2026) surfaces a distinct concern from privacy/extraction attacks: can you trust what memory itself contains and does to reasoning — corrupted consolidation, sycophantic over-reliance on stored user views, and stale/current facts silently coexisting in the same bank?
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| When Memory Becomes Authority: Benchmarking Authority Collapse at the Consolidation Boundary (AuthMem-Bench) | Zhan, Zhang, Guo, Zhao, Liu | Aug 2026 | Consolidation imposes an implicit authorization boundary — it decides whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. Names authority collapse: consolidation preserves the claim while erasing the source constraints governing its authorized use, so the memory implies greater authority than its source permits. AuthMem-Bench is a controlled paired benchmark holding the focal claim and downstream task fixed while varying only source authority. Across 7 consolidators × 7 backbones, collapse appears in 48 of 49 configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%; end-to-end, automatically predicted and persisted authority labels cut the observed rate from 16.9% to 0.0% with benign task success essentially unchanged 📄. Memory must preserve not only what was learned but the authority under which it may be reused. ⚠️ Preprint. |
| Beyond Similarity: Trustworthy Memory Search for Personal AI Agents (MemGate) | Zhang, Chen, Ma, Hu, He, Zhang, Liu, Yang, Zhang, Jia | Jun 2026 | Names memory search itself as a trust boundary: a semantically-similar memory can still be contextually inappropriate, causing cross-domain leakage, sycophancy, tool-call drift, or memory-induced jailbreaks. Evaluates A-Mem, Mem0, MemOS, and OpenClaw (real-world persistent-state agent) and finds long-term memory behaves as a durable control channel, not just a utility layer. Proposes MemGate — a 9M-parameter, 35.1MB plug-in inserted between the vector store and backbone LLM that applies a query-conditioned neural gate, turning raw similarity search into task-conditioned memory admission, with no LLM modification or inference-time judge required. Reduces memory-induced threats across frameworks/backbones while preserving utility 📄. |
| TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory | Yang, Paul, Srinivasan, Kulkarni, Chappidi | Jun 2026 | Targets errors introduced by the write/revise/delete pipeline itself — omission, corruption, hallucinated content — which become persistent system-state failures once stored. A Memory Transition Verifier scores update transitions on coverage, preservation, faithfulness; preference pairs among candidate updates drive preference-guided RL to directly optimize consolidation behavior. SOTA on MemoryAgentBench, HaluMem, Mem-alpha; +12.14 F1 on HaluMem extraction; cuts omission/corruption/hallucination by 40.1% / 79.1% / 50.0% vs the strongest baseline per error type. |
| MemSyco-Bench: Benchmarking Sycophancy in Agent Memory | Xiang, Chen, Tang, Wei, Ning, Lin, Zhang, Su | Jul 2026 | Identifies memory-induced sycophancy: agents over-align with a user's stored prior views at the cost of factual accuracy or objective reasoning. Existing memory benchmarks check whether memories are stored/retrieved/updated correctly but not how they bias downstream reasoning. Five tasks probe whether an agent can reject memory as evidence, respect its applicable scope, resolve memory-vs-objective-evidence conflicts, track updates, and still personalize appropriately. First benchmark to treat memory reliance itself as a failure mode. |
| A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory | Shi, Tang, Tung | Jul 2026 | Names ghost memory: old, current, and transition facts coexist unmarked in the memory bank and get mixed during retrieval, misleading the answer model about what's true now. Argues memory should be evaluated at three decoupled levels — bank maintenance, retrieval, answer-time resolution — since final QA accuracy hides where the failure actually occurs. ATMA is a state-aware overlay that keeps superseded/transition records, builds evidence packets for the query's requested state view, and exposes current/historical/transition labels to QA. On LTP (new LoCoMo Temporal Plus benchmark), Graphiti+ATMA improves conflict accuracy +0.240 absolute; on LoCoMo, raises temporal F1 from 0.0295 to 0.1705. |
Wave 9 (July–August 2026) surfaces the window's strongest convergent signal: in roughly four weeks, multiple independent groups imported database transaction semantics wholesale into agent memory. The shared premise — a memory write is not a belief commit — is forced by the fact that stored memory now drives irreversible external actions. Staged writes, snapshot isolation, versioned heads, natural-language rollback, and cascading repair of derived records are becoming memory-layer primitives rather than database trivia.
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| MemTX: Transactional Belief Commit for Stateful Agent Memory | Li, Wang, Lu, Chen, Li, Song, Zheng, Cai | Jul 2026 | The cluster's strongest entry and the source of its thesis. In shared-memory multi-agent settings one agent's write becomes another's premise, and eventually a tool call with real side effects — yet current systems treat every accepted write as immediately actionable truth, so a polluted tool result or a teammate's half-finished note can silently drive an irreversible action. Each record carries evidence, permissions, provenance, and validity; writes are staged inside snapshot-isolated transactions admitted by a validate-and-commit pipeline; irreversible tool calls are gated on in-flight belief state; and retracting a belief triggers typed cascading repair of its derived records and tool side effects. Two invariants — action-safety gating and cascade-repair completeness — are machine-checked via property-based testing and bounded exhaustive enumeration of 5.5M protocol states, with zero violations. Across five backbones from three model families it leads all eight baselines with paired-McNemar significance on four and ties the strongest fifth, and is the only method with zero downstream harm on every backbone 📄. "Backbone capability does not substitute for commit discipline." ⚠️ Preprint. |
| MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery | Cui, Tang, Yao, Meng, Ma, Jia | Jul 2026 | A governance layer outside the answer model, supplying the transaction boundary existing systems lack: Ordered PatchTest verifies whether an update is actually supported by its source, a Temporal Resolver selects the visible version when facts conflict, and a durable snapshot journal restores application-visible state after a fault. On an item-disjoint audit it accepts all 60 supported originals and rejects all 179 hard negatives; under persistent multi-key faults on LongMemEval-S and LoCoMo states it restores the complete declared active map without knowing the actual physical write set. Highest average F1 across all twelve answer-model configurations on MemoryAgentBench FactConsolidation, beating Dense by 17.06–24.07 points in five representative settings 📄. ⚠️ Preprint. |
| ChronoMem: Version Control and Semantic Rollback for LLM Agent Memory | Su, Xu, Zuo, Bertino | Jul 2026 | Attacks forward-only evolution: memory systems continuously accumulate, consolidate, and overwrite with no principled mechanism to inspect, version, or revert — leaving agents brittle under corrections, concept drift, and corruption, particularly after they have already been exposed to subsequent information. ChronoMem commits whole-memory snapshots at every write, maintains structured version histories, and maps natural-language undo intents to concrete historical versions through hybrid lexical + semantic retrieval, rank fusion, and reranking. Introduces a post-exposure counterfactual protocol: can the agent answer queries and summarize history as if future updates had never occurred? Integrated into Google's production-ready, open-source Agent Development Kit; claims the first open-source system and benchmark for systematic semantic global memory rollback. ⚠️ Preprint. |
| TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management | Xiao, Xu, Zhang, Chen, Shi | Aug 2026 | Retires the binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification — choices that may share a label while producing fundamentally different memory states. TARL maps each statement to one of five executable actions, identifying the affected memory, resolving its temporal scope, comparing source reliability, and updating accepted, pending, and rejected ledgers. Trained by comparing the memory states produced by alternative update operations. Ships TARL-Mem, a benchmark with fine-grained action labels and next-state targets; improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems | Pollertlam, Kornsuwannawit | Aug 2026 | The cost-side counterpart to accuracy benchmarking: Mem0, Hindsight, and Mastra Observational Memory against two reference strategies (fixed-size rolling window, full-transcript resubmission), two backbones, conversations up to 400 turns, every cost measurement paired with accuracy on 665 LoCoMo questions. Three findings, all awkward for the field: (1) serving cost cannot be predicted from conversation length and message size alone — a regression that tracks the reference strategies closely misses the memory systems by 18–69%, cost being driven by internal memory behavior; (2) break-even against full transcript is highly sensitive to system and backbone, ranging from the first tens of turns for the cheapest to never within 400 turns for the most expensive; (3) no system wins on both axes — accuracy spans 21–54%, and backbone choice drives cost as much as the memory system does 📄. ⚠️ Preprint. |
| Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents | Alake, Bernardis, Cayet, Engel, Hilloulin, Hong, Hosler, Kavantzas, Kossyk, Le, Patra, Talamadupula, Venzin (Oracle) | Jul 2026 | Database-native memory substrate built on Oracle Database, framed as a full lifecycle: ingestion → extraction → consolidation → retrieval → summarization → revision/removal, with a layered active-core / passive-store split and explicit scope control across users, agents, threads. Reports 93.8% LongMemEval accuracy using ~10.7x fewer tokens than flat-history baselines. Notable as a large enterprise-vendor entry (13 authors, Oracle) validating the memory-as-lifecycle framing that most of this list converges on independently 📄. |
| Beyond the Context Window: Cost-Performance Analysis | Pollertlam, Kornsuwannawit | Sep 2025 | Compares Mem0-style fact-based memory vs long-context LLMs on LongMemEval, LoCoMo, PersonaMemv2. Break-even: memory system becomes cheaper after ~10 interaction turns at 100K context. Long-context wins on factual recall but memory is competitive on reasoning. |
| Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads | Omri, Gan, Broveak, Geens, He, Pentland, Verhelst, Weissman, Tambe (Stanford / MIT / KU Leuven) | Jun 2026 | The first systems-level characterization of agent memory: (1) a system-oriented taxonomy along four axes; (2) a phase-aware profiling harness attributing cost to construction, retrieval, and generation; (3) characterizes 10 representative systems across two benchmark suites, showing how design choices shift cost across write vs read paths; (4) derives 10 system recommendations (construction scheduling, capability floors, amortization via query volume, freshness-latency tradeoffs, fleet-scale management). The production "memory economics" lens. 📄 ⚠️ Preprint. |
Most agent memory research ignores 50 years of neuroscience. This section bridges that gap — presenting the biological mechanisms that may underlie what LLM memory papers have independently re-discovered (forgetting, gating, dual-phase encoding), alongside the spiking neural network implementations that model them directly. Cross-references to LLM papers below are 🔬 editor's synthesis.
| Paper / Project | Authors | Date | Key Contribution |
|---|---|---|---|
| A Bio-realistic Synthetic Hippocampus for Robotic Cognition | Talanov et al. | Oct 2025 | Synthetic hippocampal architecture with dual-phase operation: online sensorimotor encoding during wake, offline consolidation via SWR-triggered replay during sleep. SNN on neuromorphic substrates (≤1W). Goal-prioritised plasticity prevents catastrophic forgetting. Models the biological mechanisms that parallel what SleepGate and CraniMem implement in software 🔬. BioNanoScience. Open access. |
| The Memristive Implementation of the Hippocampus: A Hypothesis | Talanov et al. | Aug 2025 | Hardware-level hippocampal memory using stochastic polycrystalline nano-fiber mesh as memristive substrate. Key insight: the inherent randomness of material structure mimics probabilistic biological synaptic networks. Demonstrates resistive switching tunability for implementing dynamic memory functions — bidirectional replay, synaptic up/down-scaling, consolidation. BioNanoScience. Open access. |
| Simulation of Serotonin Mechanisms in NEUCOGAR Cognitive Architecture | Talanov, Gafarov, Vallverdú et al. | 2018 | Maps neuromodulatory mechanisms to computational models: dopamine → attention, serotonin → inhibition. The "cube of emotions" model. Demonstrates that mammalian emotional-state control via monoamine neurotransmitters can be re-implemented computationally. Foundation for neuromodulator-gated memory admission. Procedia Computer Science. |
| tinyHippo | Talanov | Active | CA1 + CA3 hippocampal microcircuit simulation in NEST with Izhikevich neurons. Implements bidirectional replay, theta-modulated encoding/retrieval phase separation, and SWR-triggered consolidation. The biological reference implementation — validates the mechanisms that Membrain engineers and SleepGate abstracts. MIT license. |
| Membrain | Fatykhov | Active | Neuromorphic memory bridge using FlyHash encoding and BiCameralMemory SNN (Nengo/Voja learning). Hopfield-style attractor dynamics for pattern completion (tested up to 20% noise in PoC). Stochastic consolidation with SleepSignal. gRPC API for agent integration. The engineering abstraction layer between biological models (tinyHippo) and cognitive agents (Nous). |
The Stack: These projects form a natural hierarchy — tinyHippo (biological model, validates mechanisms) → Membrain (engineering abstraction, SNN service) → Nous (cognitive agent, consumes memory). The papers provide the theoretical foundation; the repos provide working implementations.
Why this matters for agent memory 🔬: The LLM papers in this list independently converge on mechanisms that neuroscience has studied for decades. The following correspondences are editor's synthesis — the LLM papers don't explicitly cite these neuroscience sources, but the structural parallels are striking:
Papers describe mechanisms; these repos run them. Inclusion bar: (1) implements a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance; (2) actively maintained and (3) an OSI-approved license — both enforced for the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where a reader may actually vendor the code; author code for papers and benchmark/dataset repos is held to a different standard — reproduction value — so it may carry restrictive or absent license terms and may go quiet after publication, with both facts flagged inline ⚠️; (4) and either backs a paper in this list, has real production adoption, or fills a mechanism niche nothing else covers. Agent harnesses that merely bundle a vector store, general-purpose vector databases, and repos whose only claim is a self-reported benchmark ranking are deliberately excluded. Stars, license, and last-push verified via the GitHub API on 23 Aug 2026 — a snapshot for triage, not a ranking.
| Project | Stars | License | What it actually implements |
|---|---|---|---|
| mem0 | 63.8k | Apache-2.0 | Extraction-based fact memory: an LLM pass distills durable facts from raw dialogue, then reconciles them against the existing store via add / update / delete / noop. The most widely deployed reference for the extract-then-reconcile write path. Paper: Mem0. |
| Graphiti (Zep) | 30.2k | Apache-2.0 | Bi-temporal knowledge graph: edges carry validity intervals and are invalidated, not overwritten, so superseded facts remain queryable with time bounds. The closest production analogue to Memory Transactions, Versioning & Recovery. |
| cognee | 30.2k | Apache-2.0 | ECL (extract → cognify → load) pipeline turning documents and conversations into a typed knowledge graph plus vector index, with explicit prune/forget operations over the graph rather than append-only growth. |
| Letta | 24.4k | Apache-2.0 | The MemGPT lineage: OS-style paged memory where the agent edits its own core memory block through tool calls and pages the remainder to archival storage. Self-editing memory as the mechanism. Paper: MemGPT. |
| Hindsight | 20.9k | MIT | Consolidation as an explicit policy layer over four levers — importance, merge, decay, eviction — instead of an unbounded append log. Industry framing already cited in Adjacent Research; paper: arXiv:2512.12818. |
| MemOS | 10.9k | Apache-2.0 | "Memory OS": a MemCube abstraction unifying plaintext, activation (KV), and parameter memory, with scheduling and cross-task reuse between them. Rare in treating KV-cache and long-term store as one substrate. |
| Honcho | 6.8k | AGPL-3.0 | Memory as a representation of a person, not a transcript: background reasoning maintains per-peer models, queried through a dialectic API rather than similarity search. |
| MemMachine | 3.2k | Apache-2.0 | Clean two-tier service — episodic memory plus a durable user profile — with pluggable stores and a server/client split. A readable reference implementation if you are building your own. |
| MemoryOS | 1.6k | Apache-2.0 | EMNLP 2025 Oral ✅. Short / mid / long-term stores with heat-based promotion and eviction between tiers — one of the few systems whose admission and consolidation policy is inspectable code rather than a prompt. |
| redis/agent-memory-server | 308 | Apache-2.0 | Explicit working-memory (token-bounded, session-scoped) vs long-term memory split, with automatic extraction at the boundary. MCP + REST, Redis-native. |
Reference implementations for papers already indexed in this list — verified to exist and to correspond to the cited work.
| Paper (in this list) | Code | Notes |
|---|---|---|
| xMemory | HU-xiaobai/xMemory | ⭐120 · MIT. Author release for Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation. |
| AgeMem | y1y5/AgeMem | ⭐37 · ACL 2026 SAC Highlight ✅. Unified long/short-term memory management learned as tool operations. |
| CraniMem | PearlMody05/CraniMEM | ⭐7 · gated-and-bounded memory implementation; project page. ⚠️ No license file — all rights reserved by default; clear reuse with the authors. |
| HeLa-Mem | ReinerBRO/HeLa-Mem | ⭐16 · ACL 2026 Main ✅. Hebbian associative distillation, reproduces the LongMemEval-S numbers. ⚠️ No license file — all rights reserved by default. |
| StructMemEval | yandex-research/StructMemEval | ⭐11 · Apache-2.0. Raw benchmark data + supplementary code from Yandex Research; authors label it work in progress. |
| memorywire | mthamil107/memorywire | ⭐2 · Apache-2.0. Reference implementation of the wire format, including the provenance-based purge/quarantine path for poisoned stores. |
| A-MEM | agiresearch/A-mem | ⭐1.2k · MIT. Zettelkasten-style memory that links and evolves notes on write. Paper reproduction lives in WujiangXu/A-mem (NeurIPS 2025 ✅). Last push Dec 2025 ⚠️. |
Single-binary or file-backed systems, usually MCP servers, where the store is inspectable by a human and runs without a cloud dependency.
| Project | Stars | License | What it actually implements |
|---|---|---|---|
| agentmemory | 27.3k | Apache-2.0 | Full lifecycle memory for coding agents: confidence scoring, decay and consolidation, knowledge graph, hybrid search, ~20 agent adapters. ⚠️ Headline ranking claims are self-reported. |
| Engram | 6.1k | MIT | Agent-agnostic Go binary — SQLite + FTS5, MCP/HTTP/CLI/TUI, explicit conflict handling. Deterministic lexical recall as a deliberate counter-position to embedding everything (cf. ReFind in Retrieval & Recall). |
| memsearch (Zilliz) | 2.5k | MIT | Markdown files as the source of truth, Milvus for hybrid retrieval, typed episodic/procedural memory. Human-readable store = auditable memory. |
| Nocturne Memory | 1.3k | MIT | Rollbackable, visually inspectable graph memory server for MCP agents — edits are reversible operations, not blind writes. The practical counterpart to MemTX/ChronoMem-style recovery. |
| deja-vu | 1.0k | MIT | No write path: indexes the session transcripts 34 coding agents already write locally (lexical by default; optional deja embed may send text to an external embedding API) and serves them to any of the others — including pre-install history, so no cold start. Read-side mechanisms: outcome-aware ranking (a session that ended in a fix outranks one that gave up), redaction at ingest, per-project trust policy, forget/unforget. LoCoMo/LongMemEval drivers in-repo. |
| Vestige | 608 | AGPL-3.0 | Local-first Rust MCP server combining FSRS-6 decay, spreading activation, contradiction inspection, active suppression, Receipt Lock, and an inspectable dashboard. Built for causal recall — the quiet change that broke today, not the lookalike. |
| Mnemon | 507 | Apache-2.0 | Single Go binary, four-graph knowledge store with intent-aware recall, importance decay, and automatic deduplication. No API keys, one setup command. |
| Caura (ex-MemClaw) | 439 | Apache-2.0 | Governed shared memory for agent fleets: trust tiers, keystone policies, audit trails, multi-tenant scoping. The implementation counterpart to Memory Trust & Integrity and Multi-Agent Memory. |
| mengram | 189 | Apache-2.0 | Typed semantic / episodic / procedural memory where procedures are derived from recorded failures — one of the few OSS systems implementing the failure → procedure loop from Skill & Procedural Memory. |
| causal-memory | 66 | Apache-2.0 | Local-first Rust MCP server (also CLI + PyO3 bindings) where the store is one SQLite file holding facts, temporal state, and typed decision→outcome edges (caused / enabled / prevented). Spreading activation runs both ways — prevented edges carry negative activation, a GABA analogue — alongside Hebbian co-occurrence, Q-value dynamics, write-time gating, and immutable SWR-style consolidation. Reproducible benchmark harnesses (CausalEval, LoCoMo, LongMemEval) live in-repo; Hermes Agent provider plugin included. |
| Tree Ring Memory | 13 | MIT | Project-scoped Rust CLI lifecycle layer: SQLite/FTS explainable recall, deterministic consolidation, forgetting, redaction, audit, framework discovery. Memory that ages instead of accumulating raw transcript. |
| kgai | 1 | MIT | Append-only, content-addressed decision log for dev teams: every change is a new event that supersedes the prior one through an explicit link with a reason, rejected paths stay queryable (temporal versioning), recall returns only decisions in force. Deterministic graph projection, lexical recall with no embeddings, and per-writer shard sync over the team's own S3 bucket (no server). Repo-supplied config is gated behind an explicit trust approval (governance). Claude Code plugin plus a Go CLI. |
| Hyperconsciousness | 1 | MIT | Rust knowledge store with signed, encrypted append-only records and scoped, expiring grants for MCP access. Delegated grants cannot widen scope, add actions, or outlive their parent. Governance at the retrieval boundary; does not sandbox processes running as the owner. ⚠️ Developer alpha; no independent security audit. |
Neuromorphic reference implementations — tinyHippo and Membrain — live in Neuromorphic & Bio-Inspired Memory.
| Project | Stars | License | What it measures |
|---|---|---|---|
| LongMemEval | 1.0k | MIT | ICLR 2025 ✅. 500 questions across five long-term memory abilities over long interactive histories. The de-facto standard — and the number most vendor claims cite. |
| LongMemEval-V2 | 132 | Apache-2.0 | Official successor, released 2026, addressing saturation and contamination in the original. |
| LoCoMo | 1.1k | see repo ⚠️ | Very long-term conversation benchmark (ACL 2024 ✅). Widely used dataset; frozen ⚠️ — last push Aug 2024, cite the data, don't expect maintenance. |
| MemoryData | 139 | none ⚠️ | Unified harness — 4 benchmark families, 22 method presets, one runtime — built specifically to make heterogeneous memory systems comparable. Directly attacks the fragmentation problem Evaluation & Benchmarks documents. |
| GateMem | 197 | MIT | Memory governance under multiple principals sharing one store: utility, access control, and active forgetting measured together instead of accuracy alone. |
| HaluMem | 155 | CC-BY-NC-ND-4.0 ⚠️ | Operation-level hallucination benchmark: scores extraction, updating, and answering separately, exposing errors that end-to-end QA accuracy hides. License is declared by README badge only — no LICENSE file in the repo; NC-ND terms forbid commercial use and derivatives, so treat it as read-only unless you clear it with the authors. |
| Mem-Gallery | 103 | MIT | ACL 2026 Main ✅. Multimodal long-term conversational memory for MLLM agents. Last push Jan 2026 ⚠️. |
| STATE-Bench | 77 | MIT | Microsoft. Memory-agnostic enterprise workflows measuring whether an agent improves with experience, not whether it recalls a fact. Blog write-up in Adjacent Research. |
| Resource | Stars | Why it's here |
|---|---|---|
| Agent Memory Techniques | 926 | 30 runnable notebooks covering buffers, vector stores, KGs, episodic/semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, and LoCoMo evaluation. The best hands-on on-ramp before reading the papers above. |
| Awesome-AI-Memory (IAAR Shanghai) | 1.2k | Broad bilingual knowledge base spanning LLM memory and agent memory, including engineering frameworks and applications. |
| Awesome-Agent-Memory (TeleAI) | 596 | Systems + benchmarks + papers for LLM and MLLM memory — stronger multimodal coverage than this list. |
| Awesome-GraphMemory | 339 | Survey-grade depth on graph-based agent memory specifically. |
| Awesome-Agent-Memory-Papers | 229 | Paper-only index with a browsable website. |
| OpenDataBox/awesome-agent-memory | 54 | Same name, different list — a four-axis paper taxonomy, companion to the MemoryData harness above. |
These papers and projects aren't solely about agent memory but contribute relevant architectural ideas:
| Paper / Project | Date | Relevance |
|---|---|---|
| MemTools: A Unified Research Framework for Interoperable Agent Memory — Zhao, Chen, Liang, He, Wang, Zhao, Liu | Jul 2026 | Targets architectural fragmentation as the thing blocking systematic memory research: implementations couple lifecycle stages together, entangle evaluation logic with specific datasets, and barely support heterogeneous memory types. MemTools standardizes the memory lifecycle through declarative data contracts, enabling interchangeable assembly of components across systems, and orthogonally separates benchmark datasets from execution protocols so evaluation can be reconfigured independently of data. Adds a unified computational interface for coordinating symbolic, neural, and multimodal representations in one runtime. ⚠️ Authors label it work in progress; the evaluation demonstrates isolation of design variables rather than task-level performance. |
| MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories — Lovén, Sauvola, Riekki, Tarkoma | Aug 2026 | Agents share transport and can call each other's tools, but have no protocol for reconciling what they know. MELD admits every incoming claim through a five-outcome procedure (insert, merge, relate, conflict, reject) decided from three signals — scoped claim-key identity, embedding similarity, and a natural-language- inference verdict — under context and freshness gates, acting through exactly one auditable, authenticated Patch as the sole object that mutates state, with a per-claim status CRDT over standard pub/sub. The design stance is the notable part: MELD does not adjudicate truth — a detected contradiction is preserved for later adjudication, never silently resolved. Merge classifier separates at AUC 0.968 with a 0.013 false-merge rate; the status CRDT reconverges in 30/30 real partition-heal trials where last-writer-wins manages 11/30; recall-non-inferior to a centralized store at ~11% less live storage; ~3× fewer messages at matched recall. Evaluated on a real continuum spanning operator-grade 5G edge, national HPC, and a local tier 📄. ⚠️ Preprint; code and data on Zenodo. |
| memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations — Munirathinam | May 2026 (rev Jul 2026) | JSON-Schema 2020-12 wire format for 5 memory ops (remember/recall/forget/merge/expire) × 4 memory types over mem0/Letta/Cognee/Zep/pgvector backends, with an optional HITL governance channel. Not a new algorithm — a packaging of RRF/FSM/consolidation into a vendor-neutral protocol meant to compose with MCP. v3 adds a poison-recovery eval (PurgeBench) showing its provenance field is the strongest lever for recovering a poisoned store — relevant complement to TMA-NM/GhostWriter above. |
| EverMemOS — Liu, Bai, Chen et al. | Jan 2026 | Self-Organizing Memory Operating System for structured long-horizon reasoning. |
| AriGraph — Anokhin, Radionov et al. | IJCAI 2025 ✅ | Knowledge Graph World Models with episodic + semantic memory. Proceedings |
| Memory Matters More Than You Think — Zhou, Li et al. (Renmin U.) | Jan 2026 | Event-centric memory with Logic Map for agent searching and reasoning. |
| The Consolidation Problem in Agent Memory — Vectorize (Hindsight) | May 2026 | Industry framing of consolidation as a policy layer with four levers — importance, merge, decay, eviction — benchmarked across Mem0, Zep, Letta, LangChain, Hindsight. |
| Introducing STATE-Bench — Microsoft (repo) | May 2026 | Open-source, memory-agnostic benchmark measuring whether agents improve with experience on stateful enterprise tasks (travel, support, shopping) — not just recall. |
| State of AI Agent Memory 2026 — Mem0 | 2026 | Industry survey of benchmarks, architectures, and six open problems (temporal abstraction, cross-session structure/identity, app-level eval, privacy/consent, staleness). |
These papers, taken together, converge on several architectural principles:
Papers: A-MAC, CraniMem, SleepGate, ACC
The strongest systems filter before storing. CraniMem uses goal-conditioned cosine gating; A-MAC scores across 5 dimensions (utility, confidence, novelty, recency, type). The alternative — storing everything and retrieving selectively — leads to noise accumulation and proactive interference.
Papers: xMemory, SYNAPSE, AdaMem, CraniMem, Mem0, Episodic Memory
Every serious architecture separates memory by type (episodic, semantic, procedural, working). xMemory's 4-level hierarchy shows 44.91% of memories get retroactively reassigned — proving flat structures fundamentally fail. The human brain doesn't use one store; neither should your agent.
Papers: SleepGate, CraniMem, SYNAPSE, A-MEM
Raw episodic traces should consolidate into durable semantic knowledge over time. SleepGate's sleep cycles, CraniMem's replay-based promotion, and SYNAPSE's spreading activation all implement variants of this. Without consolidation, memory becomes a write-only log.
Papers: SleepGate, Procedural Memory, xMemory
SleepGate reduces proactive interference from O(n) to O(log n) through learned forgetting. This is the most under-implemented capability in agent frameworks — most systems only add memories, never remove or decay them. Forgetting is what keeps memory useful rather than just large.
Papers: xMemory, Memory Survey, Anatomy of Agentic Memory
The clearest thesis across this literature: retrieval-augmented generation is a retrieval strategy, not a memory system. Agent memory requires admission, evolution, consolidation, forgetting, typed storage, and context-aware retrieval. RAG addresses only the last item.
Papers: Anatomy of Agentic Memory, StructMemEval
Current benchmarks are saturated, metrics misalign with semantic utility, and results vary wildly by backbone model. StructMemEval reveals agents can't organize memory autonomously. The field needs better evaluation before it can claim progress.
Papers: AgeMem, EMPO², MemFactory, Memory-R1
The dominant shift in 2026: from hand-coded admission/retrieval rules to reinforcement-learned memory policies. AgeMem exposes all memory operations as tool-based actions and learns when to use them via GRPO. EMPO² takes this further by using memory for exploration, not just recall. MemFactory provides the infrastructure to train these policies. This is the most consequential paradigm shift since the field recognized that RAG ≠ memory.
Papers: ADAM, FSFM
April 2026 surfaced the first systematic memory-extraction attack (ADAM, up to 100% ASR) — and the first forgetting frameworks (FSFM) that explicitly frame selective forgetting as a privacy primitive, not just a retention/efficiency one. Expect memory threat-modeling, audit interfaces, and unlearning APIs to follow.
Papers: Experience Compression Spectrum, Externalization in LLM Agents
The memory and skill-discovery communities have been solving the same problem (extract reusable knowledge from interaction traces) without citing each other — cross-community citation rate is below 1%. The Compression Spectrum reframes both as points on a single axis (5–20× → 50–500× → 1000×+ compression), and identifies the missing diagonal: no current system supports adaptive cross-level compression.
Papers: MemReader (active extraction), SCG-MEM (schema-constrained generation), HeLa-Mem (Hebbian distillation)
Two paradigm shifts in April 2026: (1) MemReader moves extraction from passive transcription to an RL-learned WRITE/DEFER/RETRIEVE-CONTEXT/DISCARD policy. (2) SCG-MEM replaces dense retrieval with schema-constrained decoding (Piaget-style assimilation vs accommodation). Together they suggest the future of memory is generative-constructive, not retrieve-and-paste.
Papers: MRAgent, MemPro, DeltaMem
Wave 6's headline shift: stop treating memory access as a static retrieve-then-reason step. MRAgent folds LLM reasoning into graph traversal (Cue-Tag-Content + active reconstruction), iteratively pruning paths on accumulated evidence. MemPro generalizes the idea to the whole pipeline — the memory system itself becomes an evolvable program that rewrites its own construction/retrieval logic. DeltaMem reconstructs each experience on demand by composing a root-to-leaf residual chain rather than storing flat copies. Memory is increasingly constructed at access time, not fetched.
Papers: Agent Memory Characterization, Beyond the Context Window, Personalize-then-Store
For two years memory papers reported accuracy; few reported cost. Wave 6 changes that. The first systems characterization (Omri et al.) builds a phase-aware cost profiler and a 4-axis taxonomy, deriving 10 deployment recommendations across the write and read paths. Personalize-then-Store reframes admission as a budget problem — gating transient sessions to spend the memory budget where it pays off. The question is shifting from "does memory help?" to "what does memory cost, and where does it amortize?"
Papers: TRUSTMEM, MemSyco-Bench, A-TMA
Wave 6 established memory as an attack surface (ADAM). Wave 7 shows a second, orthogonal failure class: memory that's technically retrieved correctly but still misleads. TRUSTMEM targets self-inflicted corruption during consolidation (omission, hallucination). MemSyco-Bench targets over-reliance — agents favoring a user's stored prior view over current evidence. A-TMA targets ghost memory — stale and current facts silently coexisting in the same bank. None of these are extraction attacks; they're reliability failures baked into ordinary write/retrieve operations.
Papers: MemProbe, Are We Ready For An Agent-Native Memory System?
MemProbe's headline finding — task completion nearly saturates even for a memoryless baseline, while structured state recovery from the memory artifact itself stays around 0.6 — is a direct rebuke of end-to-end accuracy as a memory metric. "Are We Ready" pushes the same point systemically: decomposing 12 memory systems into 4 modules shows no architecture dominates, and effectiveness is bottleneck-specific, not a single leaderboard number. Both extend Anatomy of Agentic Memory's Feb-2026 critique from theory into large-scale measurement.
Papers: What to Keep What to Forget (Rate–Distortion View), Is Agent Memory a Database?
Two Wave 7 papers step back from architecture-of-the-week and ask for formal foundations. The rate–distortion paper shows KV-cache eviction, prompt pruning, bounded recurrent state, and agent memory consolidation are the same budget-constrained retain/discard decision, and that all four fail identically (irreversible, pre-query, recency/attention-biased). "Is Agent Memory a Database?" argues memory correctness is a property of the state trajectory, not individual records — no record-level storage system can satisfy it, motivating the GEM formalization. Both suggest the field's next gains come from theory, not another heuristic gate.
Papers: AuthMem-Bench, PPMF
Two independent papers landed the same finding within days, from opposite directions. Summarization is lossy in a way nobody was measuring: it preserves the claim and drops the permission. AuthMem-Bench measures it as an accident — authority collapse in 48 of 49 configurations, a 50.3% mean unauthorized-action rate for collapsed memories. PPMF weaponizes it — laundering an untrusted observation into apparent user history, preserving the trigger while erasing the low-trust source, up to 1.000 ASR. Both fixes are the same shape: persist authority as first-class metadata and match action risk against it at tool-call time. This is distinct from insight #13, which asks whether a memory is true; this asks what a memory is allowed to authorize.
Papers: MemTX, MemTxn, ChronoMem, TARL
The transactions section in one line. Agent memory is acquiring database semantics — staged writes, snapshot isolation, versioned heads, rollback, and cascading repair of derived records and tool side effects. The forcing function is that memory now drives irreversible external actions, so the cost of an unsound write is no longer a bad answer but a real-world consequence. Note the direction of travel: TARL replaces a binary Write/Hold with five actions; ChronoMem adds an undo; MemTX gates tool calls on in-flight belief state. All three treat the moment of writing as provisional by default.
Papers: ReFind, Filesystem-Based Memory, Harness the Memory, Keep It InMind
The sharpest contrarian thread this list has carried. ReFind beats HippoRAG 2 (58.2 vs 53.2) with no semantic structure at all, on a matched backbone. Filesystem-Based Memory finds organization buys retrieval economy, not accuracy — and erodes as stores grow. Harness the Memory finds no substrate dominates and that excessive retrieval actively harms sequential decision-making. Against all three, Keep It InMind argues the framing itself is wrong: the failure isn't structure versus no-structure but the query-conditioned retrieval interface — 84.0% with memory in context collapsing to ≤14.4% when the same memory must be retrieved. Read together they pose an uncomfortable question for most of this list: is the field over-investing in write-time structure to compensate for a broken read-time interface?
This list grows as the field grows. To contribute:
Papers should be peer-reviewed, at reputable workshops (e.g., ICLR MemAgents), or highly cited preprints. We prioritize papers with novel architectural ideas over incremental benchmark improvements.
For repositories (the Implementations section), the bar is different: a repo must implement a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance — and either back a paper in this list, show real adoption, or fill a niche nothing else covers. Star counts are context, not a criterion. Self-reported benchmark rankings are not evidence; a reproducible harness is. Active maintenance and an OSI-approved license are required only in the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where the code is meant to be vendored. Author code for papers and benchmark/dataset repos is judged on reproduction value instead: a license must still be disclosed, and where it is absent or non-OSI (e.g. NC-ND, or no LICENSE file at all) — or where the repo has been quiet for more than six months — the row is flagged ⚠️ rather than dropped. A frozen benchmark that everyone cites still earns its place; it just gets its last-push date stated.
This list is released under CC0. The papers themselves retain their original licenses.
A curated collection of research papers on memory systems for LLM-based agents — covering architecture, retrieval, forgetting, consolidation, evaluation, the cognitive science that inspires it all, and the neuromorphic hardware that may one day implement it.
Why this list? Every agent framework bolts on a vector store and calls it "memory." These papers show what memory actually requires: admission control, consolidation loops, forgetting mechanisms, typed multi-store architectures, and retrieval strategies that go far beyond cosine similarity.
Maintained by @tfatykhov · Built alongside Nous, a cognitive AI agent with typed multi-store memory.
| If you want to... | Start with |
|---|---|
| Understand the landscape | Memory Survey (taxonomy) → Anatomy (critical analysis) |
| Build agent memory | xMemory (why RAG isn't enough) → A-MAC (admission control) → Mem0 (production system) |
| Add forgetting/consolidation | SleepGate → CraniMem |
| Evaluate your memory system | StructMemEval → Anatomy (why benchmarks are broken) |
| Train memory with RL | AgeMem (tool-based ops) → EMPO² (exploration) → MemFactory (framework) |
| Understand memory privacy risk | ADAM (extraction attack) → FSFM (forgetting as defense) |
| Unify memory ↔ skills ↔ rules | Experience Compression Spectrum → Externalization |
| Make memory writes safe & reversible | MemTX (transactional belief commit) → ChronoMem (versioning + rollback) |
| Ship something this week | Implementations & Reference Code (curated repos, mechanism-by-mechanism) |
| Explore the neuroscience | Neuromorphic section |
| Symbol | Meaning |
|---|---|
| ✅ | Venue confirmed — verified via official proceedings, OpenReview, or arXiv metadata |
| 📄 | Self-reported — metrics from authors' own evaluation; exercise caution |
| ⚠️ | Unverified — plausible claim but not independently confirmed |
| 🔬 | Editor's synthesis — cross-paper connection identified by the maintainer, not claimed by original authors |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem | Tang | Aug 2026 | The first serious attempt to say what a memory is, formally: memory is a basis, knowledge is its span, and answerability is a coverage problem — a query is answerable exactly when some single item in the span covers it. Optimal memory becomes the capacity-constrained maximizer of expected coverage, tracing a utility–capacity frontier that serves as a common yardstick for comparing memory systems. Also treats noise (coverage vs precision when the store holds false claims) and formalizes continual memory as a sequential MDP — memory is the state, writing is the action, query-time utility is the delayed reward. Instantiated concretely on Homer's Odyssey. ⚠️ Preprint; solo-author. |
| Memory in the Age of AI Agents: A Survey | 47 authors | Dec 2025 | Unified taxonomy across 3 lenses: Forms (token/parametric/latent), Functions (factual/experiential/working), Dynamics (formation/evolution/retrieval). Distinguishes agent memory from RAG and context engineering. Hugging Face Daily Paper #1 ⚠️ (unverifiable — HF archive page returns 404). |
| Anatomy of Agentic Memory | Jiang, Li, Wei et al. | Feb 2026 | Structured taxonomy of Memory-Augmented Generation (MAG) systems. Critically analyzes empirical fragility: benchmark saturation, metric misalignment, backbone-dependent variance, overlooked latency costs. Evaluates 5 systems (LOCOMO, A-Mem, MemoryOS, Nemori, MAGMA). |
| The Landscape of Agentic RL for LLMs | Guibin Zhang et al. (Oxford, Shanghai AI Lab, NUS, UIUC, UCL) | Sep 2025 | Synthesizes 500+ works. Reframes LLMs as autonomous agents using POMDPs. Covers planning, tool use, memory, reasoning, self-improvement. |
| From Static Templates to Dynamic Runtime Graphs | Yue et al. (IBM Research) | Mar 2026 | Agentic Computation Graphs (ACGs) framework — distinguishes workflow templates, realized graphs, and execution traces. Organizes ~40 papers by when structure is determined. |
| Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers | Pengfei Du | Mar 2026 | Write–manage–read loop formalization. 3D taxonomy: temporal scope × representational substrate × control policy. Five mechanism families: context-resident compression, retrieval-augmented stores, reflective self-improvement, hierarchical virtual context, policy-learned management. Proposes vision of a "foundation model for memory control." |
| Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering | Zhou, Chai, Chen, Guo, Shan, Song, Xu et al. | Apr 2026 | Unified review across four externalized components: memory stores, reusable skills, interaction protocols, runtime harness. Argues modern agent capability comes from reorganizing the runtime around the model, not from changing weights. Connects memory literature with skill-discovery and protocol-design literatures that rarely cite each other. |
| Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents | Zhang, Wang, Cui, Qiu, Li, Zhu, He | Apr 2026 | Positions memory/skills/rules on a single compression axis (5–20× episodic, 50–500× procedural, 1000×+ declarative). Citation analysis of 1,136 references finds <1% cross-community citation between memory and skills literatures. Identifies the "missing diagonal" — no current system supports adaptive cross-level compression. |
| From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms | Luo, Tian, Cao, Luo et al. | May 2026 | ACL 2026 Findings ✅ — Maps the field's evolution from passive storage stores toward experience-centric, self-evolving memory. Companion paper-list: Evolving-LLM-Agent-Memory-Survey. |
| Graph-based Agent Memory: Taxonomy, Techniques, and Applications | Yang, Zhou, Xiao, Dong et al. | Feb 2026 | First focused taxonomy of graph-based agent memory: node/edge designs, construction strategies (entity-centric vs event-centric vs hybrid), retrieval (one-hop vs multi-hop vs subgraph), and update/forget operators. Useful companion to A-MEM, HeLa-Mem, GAM. |
| Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory | Orogat, Mansour | May 2026 | Argues record-level correctness (rows, embeddings, edges) cannot satisfy long-term memory's needs, causing four recurring failure modes: unregulated growth, missing semantic revision, capacity-driven forgetting, read-only retrieval. Formalizes Governed Evolving Memory (GEM) — state-level operators (ingestion, revision, forgetting, retrieval) with six correctness conditions on the trajectory, not individual records. Prototype MemState on a property-graph backend. Proves no record-level system can satisfy the conditions regardless of storage model. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Metis: Memory Foundation Model | Zhang, Guo, Sun, Zhang, Hao, Lin et al. (17 authors) | Jul 2026 | Challenges this list's prevailing external-module framing. Argues agent memory should be native to the backbone, and formalizes native memory as (a) a persistent, dynamically evolving memory state inside the model and (b) native memory procedures that store and use information through model computation. Metis equips a foundation model with a native memory state accessed via memory attention, acquired through mid-training on large-scale memory-specific data. Online memory maintenance is gradient-free — a memory update requires only a forward pass, and all learned weights remain frozen at inference. Project and model checkpoints released. ⚠️ Preprint. |
| Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability | Zhou, Yu, Wei, Wu, Ouyang, Jiao et al. (11 authors) | Jul 2026 | The first systematic study of the memory design that is actually deployed by default — a directory tree of markdown files the agent itself reads, writes, and reorganizes with generic file tools — which prior research largely passed over. Formalizes it as three roles around one memory filesystem (management, search, execution agents), unifying declarative memory and skills in a single store. The default's two working assumptions do not hold up: what organization reliably buys is search economy (organized stores roughly halve retrieval cost where material is large), not better answers — no agent measured converts organization into accuracy, and organization erodes as the store grows for all but the strongest management agent. Changing the tool set alone reshapes the store as strongly as swapping the model 📄. ⚠️ Preprint. |
| MAGE: Memory as Agent-Guided Exploration | Chen, Lai, Feng, Han, Zhang, Lu, Li et al. | Jun 2026 | Argues semantic-similarity organization mismatches execution-state dependencies on long-horizon tasks — it fragments decision trajectories and mixes valid/erroneous traces. Stores interactions in a hierarchical state tree; agent state derives from the active root-to-current path (subgoal summaries + recent traces + branch hints). Four coupled ops: Grow, Compress, Maintain, Revise. +7.8–20.4pp task success on MemoryArena, -55.1% tokens 📄. |
| Mechanistic Attention Guidance for Agent Memory Refinement (AGMR) | Hong, Qiu, Wang, Yao | Jul 2026 | Existing self-evolving memory only inspects textual outputs (trajectories, reflections), risking unreliable error attribution and hallucinated edits. Uses retrieval-head attention as a mechanistic signal — aggregates attention over memory segments × decision steps into a context-utilization matrix that reveals actual memory-use patterns. AGMR corrects/enhances memory on failure, simplifies it on success, and re-executes to verify each update. Beats text-only refinement baselines on interactive decision-making benchmarks. Code released. ⚠️ Preprint. |
| Supra Cognitive Modes (SCM): A Routed Architecture for Agent Memory | Tobkin, Yang | Jul 2026 | Routes each query — factual lookup, relation-chain/current-state reasoning, or broad long-history synthesis — to a matched retrieval+synthesis payload over one shared ingest substrate (multi-granularity embeddings, extracted triples, fact-version metadata). A frozen semantic classifier + runtime gates dispatch among fused lexical/dense lookup, graph/multi-hop, and stratified long-form synthesis. Reports 84.87% LoCoMo factoid / 68.61% adversarial abstention, 61.49% MAB, 86.00% LongMemEval — but the authors themselves flag causal routing effects and efficiency gains as outside the available evidence 📄 (unusually candid self-limitation disclosure). ⚠️ Preprint. |
| A-MEM: Agentic Memory for LLM Agents | Xu, Liang, Mei, Gao, Tan, Zhang | Feb 2025 | Dynamic memory organization inspired by Zettelkasten. Creates interconnected "memory notes" with content, metadata, and explicit links. Agent autonomously decides when to create, update, or link memories. NeurIPS 2025 ✅. |
| SYNAPSE: Episodic-Semantic Memory via Spreading Activation | Jiang, Chen, Pan et al. | Jan 2026 | Memory as a dynamic graph where relevance emerges from spreading activation rather than pre-computed vector links. Lateral inhibition and temporal decay highlight relevant sub-graphs. Triple Hybrid retrieval. |
| MAGMA: Multi-Graph Agentic Memory Architecture | Jiang, Li, Li, Li | Jan 2026 | Multi-graph based architecture separating different memory concerns into distinct graph structures for AI agents. ACL 2026 Main ✅. |
| Mem0: Production-Ready AI Agents with Scalable Long-Term Memory | Chhikara, Khant, Aryan, Singh, Yadav | Apr 2025 | Scalable memory-centric architecture with dynamic extraction, consolidation, and retrieval. Graph-based variant for complex relational structures. 26% improvement over OpenAI memory, 91% lower latency, 90%+ token savings 📄 (self-reported). Evaluated on LOCOMO. |
| CraniMem: Neurocognitively Motivated Gated & Bounded Memory | Mody, Panchal, Kar, Bhowmick, Karani | Mar 2026 | Cranial-inspired dual-store: bounded FIFO episodic buffer + knowledge graph (long-term). Goal-conditioned gating. Utility tagging (Importance + Surprise + Emotion). Beats Mem0 by +57.6% on noisy HotpotQA 📄 (self-reported; "noisy" = authors' own distractor injection, not standard benchmark). ICLR 2026 MemAgents Workshop ✅. Code |
| AdaMem: Adaptive User-Centric Memory | — | Mar 2026 | Working, episodic, persona, and graph memories. Key innovation: question-conditioned retrieval planning — resolves target participant first, builds retrieval route combining semantic + relation-aware graph expansion. SOTA on LoCoMo and PERSONAMEM. |
| PlugMem: A Task-Agnostic Plugin Memory Module | Yang, Galley, Wang, Gao, Han, Zhai (Microsoft Research, UIUC) | Mar 2026 | Converts raw interactions into propositional (facts) + prescriptive (skills) knowledge, organized in a knowledge-centric memory graph. Knowledge — not entities or chunks — is the unit of memory access. Single plug-and-play module outperforms task-specific designs across 3 diverse benchmarks. Highest information density: more useful info per context token consumed 📄. |
| Memora: Harmonic Memory Representation | Xia, Zhang, Dixit, Harimurugan, Wang, Ruhle, Sim, Bansal, Rajmohan (Microsoft Research) | Feb 2026 | Proves that RAG and KG memory are special cases of this unified framework. Primary abstractions index concrete memory values; "cue anchors" expand retrieval beyond semantic similarity. New SOTA on both LoCoMo AND LongMemEval 📄. ICML 2026 ✅. |
| MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning | Lin, Zhang, Lu, Liu, Tang, He, Zhang, Wang (Microsoft Research) | Mar 2026 | Multi-agent framework: Meta-Thinker → Memory Manager → Query Reasoner. Backward path innovation: synthesizes probe QA pairs, verifies memory, converts failures into repairs BEFORE finalizing. Plug-and-play — improves 3 different storage backends on LoCoMo. Code |
| H-Mem: Hybrid Multi-Dimensional Memory | Ye, Huang, Chen, Zhang (Rutgers) | Mar 2026 | Organizes memory across time AND topic dimensions simultaneously. Mimics associative + hierarchical properties of human memory. EACL 2026 ✅. |
| H-MEM: Hierarchical Memory for High-Efficiency Long-Term Reasoning | Sun et al. | Mar 2026 | Multi-level memory storage with positional index encoding of sub-memory at each layer. Confidence-weighted retrieval — attaches memory weights to provide LLMs with uncertainty reference. Key finding: retrieval is ineffective without structured hierarchical storage. EACL 2026 ✅. |
| GAM: Hierarchical Graph-based Agentic Memory for LLM Agents | Wu, Zhang, Lin, Xu, Xu, Chen, Zou et al. | Apr 2026 | Explicitly decouples encoding from consolidation: an event-progression graph captures stream updates online; integration into the topic-associative network is deferred until a semantic shift is detected. Addresses the tension between stream-based fluidity and structured retention. |
| HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents | Zhu, Li, Zhang, Liu, Yang | Apr 2026 | ACL 2026 ✅ — Bio-inspired dual-graph: (1) episodic graph evolves via Hebbian co-activation; (2) semantic store populated via Hebbian Distillation — a Reflective Agent identifies densely-connected hubs and distills them into reusable semantic knowledge. Beats prior SOTA on LoCoMo across 4 categories with fewer context tokens 📄. Code |
| LightMem: Lightweight LLM Agent Memory with Small Language Models | Zhang, Zhang, Chen, Huang, Zheng et al. | Apr 2026 | ACL 2026 ✅ — SLM-driven memory with strict online/offline separation. STM/MTM/LTM tiers with two-stage retrieval (vector coarse → semantic re-rank). ~2.5 F1 over A-MEM on LoCoMo, 83ms retrieval, 581ms end-to-end 📄. Shows that careful SLM use can replace repeated large-model memory calls. |
| MemMachine: Ground-Truth-Preserving Memory for Personalized AI Agents | Wang, Yu, Love, Zhang, Wong, Scargall, Fan et al. | Apr 2026 | Open-source system integrating short-term, long-term episodic, and profile memory in a ground-truth-preserving pipeline. Targets multi-session degradation in standard RAG. |
| Omni-SimpleMem: Autoresearch-Guided Lifelong Multimodal Memory | Liu, Ling, Qiu, Liu, Han, Xia, Tu et al. | Apr 2026 | Uses autonomous research-agent search over the design space (architecture × retrieval × prompts × data pipeline) to discover effective lifelong multimodal memory configurations. The first paper to treat memory architecture itself as a search target. |
| Human-Inspired Memory Architecture for LLM Agents | Kerestecioglu, Robsky, Vasters, Sharma, Kesselman (Microsoft) | May 2026 | Biologically-grounded architecture with six cognitive mechanisms: sleep-phase consolidation, interference-based forgetting, engram maturation, reconsolidation on retrieval, entity knowledge graphs, hybrid multi-cue retrieval. Introduces a synthetic calibration methodology that derives all thresholds without benchmark exposure — eliminates a common eval-leakage source. First streaming M-tier LongMemEval eval (475 sessions). Dedup-based consolidation: 97.2% retention precision, 58% store reduction (+21.8 pp) on a 13K-issue VSCode dataset 📄. |
| MRAgent: Memory is Reconstructed, Not Retrieved — Graph Memory for LLM Agents | Ji, Li, Hooi (NUS) | Jun 2026 | ICML 2026 ✅ — Represents memory as a Cue-Tag-Content graph where associative tags bridge fine-grained cues to contents. An active reconstruction mechanism folds LLM reasoning into memory access — iteratively exploring and pruning retrieval paths on accumulated evidence, adapting retrieval to the reasoning context while avoiding combinatorial blowup. Up to 23% over strong baselines on LoCoMo + LongMemEval with substantially lower token/runtime cost 📄. The cleanest articulation of the "reconstruction ≠ retrieval" paradigm. |
| MemPro: Agentic Memory Systems as Evolvable Programs | Liu, Wang, Wu, Huang, Tao, Song, Zhou, He (ECNU) | May 2026 | Treats the entire memory construction-retrieval (MCR) pipeline as an evolvable program, not just the memory bank. Maintains a version tree of runnable memory-system implementations; an Evolving Agent selects promising versions, diagnoses recurring failures, and synthesizes improved children via failure-mode-guided edit-debug refinement. Beats static and prompt-level evolving baselines on LongMemEval / LoCoMo / HotpotQA / NarrativeQA within a few iterations, with favorable performance-cost trade-off 📄. Code available. ⚠️ Preprint; venue unconfirmed. |
| MemDreamer: Hierarchical Graph Memory with Agentic Retrieval for Long Video | — | Jun 2026 | Extends agent memory to the multimodal long-video setting: a hierarchical graph memory plus agentic retrieval over hours-long video, exploiting the tree+graph structure for question answering. ⚠️ Preprint; venue unconfirmed. |
| H-Mem (Yu & Fang): Evolving & Retrieving Agent Memory via a Hybrid Structure | Yu, Fang, Liu, Ma (CUHK-Shenzhen / Huawei) | May 2026 | Distinct from the EACL H-Mem (Rutgers) and the Sun et al. H-MEM listed above — a third system sharing the name. Hybrid tree + graph: a temporal-semantic tree lets short-term memory evolve into a summarizing long-term store, while a parallel knowledge graph captures entity relationships; retrieval exploits both. SOTA on QA across three agent-memory benchmarks 📄. ⚠️ Preprint; venue unconfirmed. |
| MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management | Liu, Wu, Liu, Zhao, Liu, Li, Zhang, Wang, Guo, Liu | Jun 2026 | Extends agent memory to long-horizon mobile GUI agents. Introduces Context-as-Action (ConAct): context management becomes a first-class action emitted by the same policy that selects UI actions, maintaining three structured fields (folded action history, folded UI state, recent step record) instead of passively appending ReAct-style transcripts. 8B model trained on the 2,956-trajectory MemGUI-3K dataset achieves best open-data 8B performance on MemGUI-Bench and generalizes OOD to MobileWorld. Code/data to be released. ⚠️ Preprint. |
| Multi-Head Recurrent Memory Agents | Li, Yeh, Li | Jul 2026 | Diagnoses why recurrent memory agents degrade at long context: decomposing performance into capture vs retention shows retention collapse is the dominant bottleneck, caused by treating memory as one monolithic text block that every update risks overwriting. Multi-Head Recurrent Memory (MHM) partitions memory into independent heads with a stage-wise select-then-update strategy — only one head updates per step, the rest are structurally shielded. The MHM-LRU instantiation is training-free and lifts RULER-HQA retention at 896K tokens from <30% to 73.96%. Architecture, not model behavior, is the lever. 📄 ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| When Not to Write Memory: Governing False Promotion from Correlated Agent Traces (GovMem) | Qi, Xu, Li | Jun 2026 | Reframes admission as a write-path governance problem: repeated observations aren't independent evidence if copied from a shared source, induced by a shared prompt, or valid only in a narrower scope. GovMem estimates dependency-aware support, retrieves counterevidence, assigns scope, and outputs promote / reject / needs-review. Cuts false promotion 0.597→0.040 (synthetic) at 0.960 recall; on a human-labeled real-trace subset, false promotion falls to 0.032 but held-out false promotion stays 0.111 at 0.692 review burden. On 133 high-impact external coding-agent candidates, none were judged safe for automatic promotion. Positioned as a diagnostic governance design point, not a validated auto-writer — a sobering result for anyone auto-promoting agent traces into memory 📄. Accepted at MLISE 2026 ✅. |
| A-MAC: Adaptive Memory Admission Control | Zhang et al. (Workday AI) | Sep 2025 | 5-dimension scoring: Utility (LLM call), Confidence (ROUGE-L grounding), Novelty (1 - max cosine similarity), Recency, Type Prior. Learned weights vs fixed heuristics. |
| ACC: Agent Cognitive Compressor | Bousetouane | Jan 2026 | Bio-inspired memory controller replacing transcript replay with bounded internal state updated online. Compresses agent cognitive state without losing decision-relevant context. |
| MemReader: From Passive to Active Extraction for Long-Term Agent Memory | Kang, Li, Chen, Tang, Xiong, Li | Apr 2026 | First RL-trained active extraction policy (vs passive transcription). MemReader-4B uses GRPO + ReAct to evaluate value/ambiguity/completeness, then chooses WRITE / DEFER / RETRIEVE-CONTEXT / DISCARD before admission. Integrated into MemOS. Addresses memory pollution from noisy dialogue and cross-turn dependencies. |
| Personalize-then-Store: Benchmarking & Learning Personalized Memory for Long-horizon Agents | In, Kim, Park, Yoon, Park (KAIST) | May 2026 | Argues universal static storage policies waste budget on transient sessions while dropping critical long-horizon context. Introduces PerMemBench (first benchmark for personalized memory policies, with multi-year multi-domain personas) and session-level storage gating — a lightweight per-user admission filter that bypasses memory ops for transient sessions. Personalization yields large retention gains under perfect gating; accurate gating remains the open challenge 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory (ReFind) | Li, Zhang, Xu, Du, Fu, Chen | Aug 2026 | Asks how much of structured memory's benefit comes from the structure versus from competent retrieval over the raw history — and answers uncomfortably. ReFind builds no semantic structure at all: it leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls (session-aware rank fusion, local context expansion, temporal narrowing, skipping already-inspected sessions). Across ~2,800 questions under MemoryAgentBench's incremental multi-turn setting it attains the highest mean accuracy (58.2) of any system compared — above the strongest graph- and tree-based systems (HippoRAG 2, 53.2) — all under a GPT-4o-mini backbone matched to every reused baseline. Reaches 93.2±3.3 / 89.3±6.0 on LongMemEval-S/M with GPT-5-mini, with no LLM-based index construction at all 📄. ⚠️ Preprint. |
| xMemory: Beyond RAG for Agent Memory | Hu, Zhu, Yan, He, Gui | Feb 2026 | 4-level memory hierarchy. Key thesis: RAG ≠ agent memory. Submodular diversity-aware retrieval (MMR) replacing naive top-k. Uncertainty-gated adaptive expansion. Theme clustering. 44.91% retroactive reassignment rate proves flat structures fail. ⚠️ Conference status unverified. |
| SuperLocalMemory V3 | — | Mar 2026 | First information-geometric foundations for agent memory. Fisher information metric replaces cosine similarity. Riemannian Langevin dynamics for retrieval. Theoretical grounding for memory operations. |
| ExpRAG: Retrieval-Augmented LLM Agents Learning to Learn from Experience | Ferraz, Deffayet, Nikoulina, Déjean, Clinchant | Mar 2026 | Trains agents to USE retrieved trajectories in-context (retrieval-augmented fine-tuning). Standard LoRA collapses on OOD tasks; ExpRAG-LoRA generalizes to held-out hard tasks. Combines experience retrieval with fine-tuning — neither alone is sufficient. |
| To Know is to Construct: Schema-Constrained Generation for Agent Memory (SCG-MEM) | Zheng, Song, Li, Yang | Apr 2026 | Replaces dense retrieval with schema-constrained decoding. Piaget-inspired: assimilation (grounding into existing schemas) vs accommodation (expanding schemas with novel concepts). Solves "structural hallucination" — LLMs generating references to non-existent memory keys. Constructivist alternative to similarity-based recall. |
| SAM: State-Adaptive Memory for Long-Horizon Agents | — | May 2026 | Targets retrieval of scattered information over long horizons — adapting memory access to evolving task state rather than fixed top-k recall. ⚠️ Preprint; venue unconfirmed. |
| RaMem: Contextual Reinstatement for Long-term Agentic Memory | Yang, Kan, Li, Li, Qin, Li, Bogdan, Thomason | Jun 2026 | Names and addresses context collapse: retrieved memory fragments that share entities or user states can look equally relevant even when their surrounding episodic conditions (time, session, participants) differ, so similarity-based retrieval returns content-relevant but context-invalid evidence. RaMem runs four coordinated stages — evidence anchoring, recall-condition induction, validity-aware retrieval, context-preserved synthesis — to prioritize context-compatible memories over merely similar ones. +10% average F1 over strong baselines across several backbones on long-term memory benchmarks 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| TEPA: Revoking Stale Memories for Conflict-Robust Language Agents | Zhou, Ouyang, Zheng, Xiang | Aug 2026 | Makes validity an explicit state of memory. Observations are stored as keyed precedents; an active precedent is revoked when fresh evidence contradicts it under the same key, so retrieval draws only from current evidence while revoked history is preserved for audit rather than deleted. The headline result is a failure mode for the naive strategies: under full reversal, append-only and last-write-wins both score 0.210 — below the 0.309 of having no memory at all — while TEPA reaches 0.950 (50 seeds; reproduced under real file-backed execution at 0.203 / 0.298 / 0.950). Honest scoping: on clean MemoryAgentBench SH-6k, TEPA only matches a strong last-write-wins cache, confirming current-key replacement as the decisive op for single-hop facts; multi-hop and very-long-context settings expose retrieval-chain bottlenecks beyond fact-level validity tracking 📄. ⚠️ Preprint. |
| Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory | Kang, Liu, Kai, Liang, Tang et al. | Jul 2026 | Reframes retention-vs-consolidation as a budget-dependent operator choice, not an architectural commitment. Decomposes each operator's utility (Merge / Abstract / Rewrite) into a coverage effect on evidence that retention omits plus a signed replacement effect on raw evidence that already fits — their balance explains why the preferred action changes with relative budget pressure. Implemented as Offline Abstraction-Safety (OAS), a lightweight learner estimating action utility from pre-generation features with held-out harm calibration. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets while retention is preferable under loose ones; LoCoMo replicates the crossover at a smaller budget, consistent with its shorter evidence. Cross-note abstraction and merging generally outperform local rewriting 📄. ⚠️ Preprint. |
| Control-Plane Placement Shapes Forgetting | Yang | Jun 2026 | 13-configuration architectural study of where an LLM should sit in the memory pipeline — recall plane (extensively benchmarked) vs. control plane that mutates via supersede/release/purge (largely untested). Finding: deterministic primitives handle lexical/temporal forgetting but fail canonicalization (5% identifier-obfuscation, 0% cross-lingual); inscribe-time LLM fixes canonicalization (100%) but can't do intent-aware deletion (0%); a mutation-time hook recovers intent-aware deletion (78–85%) and lifts nearly every category simultaneously (91.7–93.2% overall) at $0.17/385-case run. Releases ForgetEval (1000-case + 385-case adversarial suite, 10-annotator Fleiss' κ=0.958) plus a 6-method Adapter Protocol. Core claim: production failures are predominantly forgetting failures, not recall failures — yet benchmarks measure only recall. Code/benchmark released (MIT) 📄. |
| SleepGate: Sleep-Inspired Forgetting for LLMs | Xie (Kennesaw State) | Mar 2026 | Learned sleep cycle over KV cache addressing proactive interference. Conflict-aware temporal tagger + forgetting gate + consolidation module. Reduces interference horizon from O(n) to O(log n). 99.5% retrieval accuracy at PI depth 5 vs 23% baseline 📄 (self-reported; extraordinary gap warrants independent replication). Supersession detection via binary σ flag. |
| SCM: Sleep-Consolidated Memory with Algorithmic Forgetting | Shinde | Apr 2026 | Five components inspired by human memory: limited-capacity working memory, multi-dimensional importance tagging, offline sleep-stage consolidation with distinct NREM and REM phases, intentional value-based forgetting, and a computational self-model for introspection. Reports perfect 10-turn recall, 90.9% noise reduction, sub-millisecond search 📄. Research preview. |
| FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory | Gu, Xiong, Wang, Ren, Li, Zhang, Guo et al. | Apr 2026 | Grounds selective forgetting in hippocampal indexing/consolidation theory and the Ebbinghaus curve. Argues that in resource-constrained environments a forgetting mechanism is as crucial as retention — and that memory security depends on the agent's ability to actively drop sensitive history. |
| Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning (LRE) | Lia, Mazumder | Jun 2026 | Reframes eviction as a fidelity problem, not a compression problem: dropping one load-bearing detail (an access token, a required path) fails the task outright. Learned Relevance Eviction (LRE) is a few-kilobyte, CPU-only, LLM-free scorer that learns which history units are load-bearing and keeps them verbatim. Matches full-history-retention accuracy while cutting peak context up to 52%; on LoCoMo gives the best budgeted answer quality reading 68% fewer tokens; annotation-free training on the system's own behavior recovers 95% of supervised effectiveness. Evidence that cheap learned relevance can replace LLM-mediated summarization for eviction 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| AgeMem: Agentic Memory — Learning Unified LTM and STM Management | Yu, Yao, Xie, Tan, Feng, Li, Wu | Jan 2026 | Exposes store/retrieve/update/summarize/discard as tool-based actions. Agent learns WHEN and WHAT to do via 3-stage progressive RL + step-wise GRPO. ACL 2026 SAC Highlight ✅ — Outperforms all heuristic-based baselines across 5 benchmarks. Key thesis: memory operations should be learned, not hard-coded 📄. |
| EMPO²: Exploratory Memory-Augmented On- and Off-Policy Optimization | Liu, Kim, Luo, Li, Yang | Feb 2026 | Uses memory not just for recall but for exploring novel states. Hybrid on/off-policy RL — 128.6% improvement over GRPO on ScienceWorld 📄. Adapts to new tasks with just a few memory-augmented trials, zero parameter updates. ICLR 2026 ✅. |
| MemFactory: Unified Inference & Training Framework for Agent Memory | Guo, Li, Tang, Xiong, Li | Mar 2026 | "LLaMA-Factory for memory agents" — first unified modular framework. Lego-like plug-and-play memory components. Natively integrates GRPO for RL-based memory policy training. Supports Memory-R1, RMM, MemAgent paradigms out of the box. Up to 14.8% improvement over base models 📄. |
| Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents | Yan, Bahloul, Nie, Schwarzmann, Trivisonno, Tresp, Ma | May 2026 | Identifies a fundamental flaw in GRPO for memory-RL: once rollouts write different memories, they no longer share the same effective environment, so trajectory-level group comparisons are unfair. Introduces LoGo-GRPO (Local + Global) — global keeps end-to-end long-horizon reward; local re-rollouts compare memory-op outcomes from the same intermediate state. Shared-parameter co-learning for fact extractor + memory manager. Progressive curriculum 8→16→32 sessions. |
| What Training Data Teaches RL Memory Agents: Curriculum Effects in Memory-Augmented QA | He, Lin, Liu, Wu, Xie, Zhou, Xiao | May 2026 | Controlled study holding architecture/RL/hyperparams fixed and varying only curriculum across in-domain (LoCoMo), mixed (LoCoMo+LongMemEval), and OOD (LongMemEval). Key findings: curriculum is a fine-grained lever on specialization, not a uniform scaling factor; per-type differences dwarf aggregate differences (single-number benchmark comparisons systematically underreport); binary EM reward produces no signal at G=4 group size — continuous rewards needed for single-GPU regime. Practical RL-memory training playbook. Code |
| SaliMory: Orchestrating Cognitive Memory for Conversational Agents | Zhang, Zhang, Jiang, …, X. L. Dong (Meta) | Jun 2026 | Trains a single LM to manage a cognitively-structured memory (user facts, preferences, working memory). Introduces a hierarchical stage-wise process reward + reward-decomposed contrastive (GRPO-style) refinement, giving isolated supervision to distinct ops (selective filtering, consolidation, cue-driven recall) end-to-end — sidestepping the credit-assignment bottleneck of standard RL over multi-stage pipelines. Cuts memory-attributed failures by ~⅓, +10% end-to-end accuracy, more than doubles the Good-Personalization rate 📄. ⚠️ Preprint. |
| JAMEL: Joint Agent Memory & Exploration Learning via Novelty Signals | Tian, Weng, Kong, …, Y. Li (Tsinghua / AIR) | Jun 2026 | Co-trains the memory module and exploration policy in a mutually-dependent loop: exploration needs memory to tell exhausted from unseen behaviors, while novelty-seeking interaction supplies annotation-free supervision (e.g., code coverage) for the latent memory. Generalizes to unseen environments; rivals a closed-source model while reducing tokens 📄. Code/model open-sourced. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| The Missing Memory Hierarchy: Demand Paging for LLM Context Windows | Mason (UBC / Georgia Tech) | Sep 2025 | Maps OS virtual memory concepts to LLM context: physical memory = context window, virtual memory = persistent state, page table = retrieval handles, page fault = re-request evicted content. Analyzed 857 sessions, 54,170 API calls, 4.45B effective input tokens. |
| What to Keep, What to Forget: A Rate–Distortion View of Memory Compaction in LLMs and Agents | Colaco, Lahjouji | Jul 2026 | Unifies four separate compaction literatures — KV-cache eviction/quantization, prompt pruning/distillation, bounded architectural state, and agent memory consolidation — as instances of one rate–distortion problem: what to retain vs discard, at what fidelity, under a resource budget, to preserve downstream utility. Builds a layer-agnostic lower bound and a seven-axis taxonomy that lets mechanisms transfer between layers that have never been connected (serving-stack KV management ↔ agent long-term memory). Finding: at every layer the keep/discard signal is attention magnitude or recency, and it fails the same way everywhere — discarding before the query is known, with no way to undo it. Proposes a cross-layer benchmark that no existing benchmark provides. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents | Huang, Zhang, Wu, Chen, Jiang, Yang et al. — incl. Ying Nian Wu, Kai-Wei Chang, Philip S. Yu, Aylin Caliskan (15 authors) | Aug 2026 | Controlled harness comparing memory substrates — the underlying medium memory is represented in: dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms — across 3 backbones × 4 benchmark suites under 26 performance and efficiency metrics. Headline: no single substrate consistently dominates. Broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context; substrates that do well at moderate history lengths become costly or brittle at longer horizons. Motivates substrate routing as a necessary component of adaptive agent memory. ⚠️ Preprint; code on acceptance. |
| Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory | Li, Du, Xu, Mao | Jul 2026 | Names an assumption so natural it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks it — a tree-nut allergy should change the answer to a macaron request through almond flour, yet the two texts share no cue a retriever can see. InMind is a 125-task, expert-verified benchmark across ten life domains (113 tasks grounded in citable public sources) whose paired controls separate three explanations existing evaluations conflate: never stored, no bridging knowledge, or stored-but-never-surfaced. The verdict is clean: with the decisive memory placed in context the backbone answers 84.0% of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic systems reach at most 14.4% — even though they recall those same facts on demand at up to 100%. An embedding with 8× the dimensionality does not close the gap. Locates the failure in the query-conditioned interface itself, naming routing — deciding which facts stay visible — as the open problem. |
| MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation | Wang | Jun 2026 | Controlled one-variable-at-a-time protocol on LongMemEval-S across 3 model families, exposing confounds that inflate architecture claims: verbatim RAG ≈ full-context GPT-4o-mini (47.2% vs 49.8%, p=0.34) but the ranking reverses by model (Gemini +14pp from full context; Sonnet +31pp from RAG, partly by refusing 63% of full-context queries); swapping only the embedding model shifts accuracy +6.2pp and flips which system wins; agent self-memory (42%) underperforms basic retrieval (47%); Mem0 matches cloud RAG on only 2/6 question types at 50x the cost. Recommends fixing embedding models across comparisons and reporting write-path cost before attributing gains to architecture — directly relevant to how Infini Memory/MAB-style comparisons should be read 📄. |
| StructMemEval: Evaluating Memory Structure in LLM Agents | Shutova, Olenina, Vinogradov, Sinitsin | Feb 2026 | Tests agents' ability to organize memory (not just retrieve). Tasks: transaction ledgers, to-do lists, trees. Key finding: LLMs don't spontaneously recognize when to apply memory structure — they succeed only when explicitly prompted. |
| MemoryAgentBench: Evaluating Memory via Incremental Multi-Turn Interactions | Hu, Wang, McAuley (UC San Diego) | Jul 2025 (revised Jun 2026) | Tests 4 memory competencies in realistic incremental accumulation. Key finding: no current method masters all 4 competencies simultaneously. ICLR 2026 ✅. Code |
| From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (Memora + FAMA) | Uddin, Shubham, Blanco, Baral, Wang (ASU) | Apr 2026 | ACL 2026 Findings ✅ — Introduces Memora, a weeks-to-months benchmark over three tasks: remembering, reasoning, recommending. Introduces FAMA (Forgetting-Aware Memory Accuracy) — a metric that penalizes reliance on obsolete or invalidated memory rather than just rewarding recall. Evaluation of 4 LLMs and 6 memory agents finds frequent reuse of invalid memories and failures to reconcile evolving knowledge. ⚠️ Not to be confused with Microsoft's Memora: Harmonic Memory Representation (2602.03315). |
| STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? | Chao, Bai, Sheng, Li, Sun | May 2026 | Benchmark for Implicit Conflict — a later observation invalidates an earlier memory without explicit negation, requiring inference + commonsense to detect. 400 expert-validated scenarios, 1,200 queries, contexts up to 150K tokens. Three-dimensional probing: State Resolution, Premise Resistance, Implicit Policy Adaptation. Best frontier LLM only 55.2% overall. Companion prototype CUPMem uses structured state consolidation + propagation-aware search. Complements FAMA. |
| MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks | He, Wang, Zhi, Hu, Chen, Yin, Wu, Ouyang, Wang, Pei, McAuley, Choi, Pentland | Feb 2026 | Memory-Agent-Environment loop benchmark across web navigation, preference-constrained planning, progressive information search, sequential formal reasoning. Key result: agents near-saturated on LoCoMo perform poorly in MemoryArena, exposing a gap between long-context memorization benchmarks and interdependent agentic memory use. Project |
| MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery | Ma, Zhou, Huang, Yang, Ma, Wang, Li, Miao, Yu, Wang | Jun 2026 | Argues memory should be evaluated as an auditable post-interaction artifact, not just through downstream task success. Simulates 50 users each carrying a hidden, taxonomy-anchored 31-dimension state bank across leak-controlled tasks, then reconstructs that state from the agent's resulting memory (full-store and top-k access) and scores it against ground truth. Key finding: task completion nearly saturates even for a memoryless baseline, while category-balanced state recovery stays moderate (~0.6) and drops further under top-k — successful assistance and recoverable memory are distinct capabilities. First benchmark to score memory recovery directly. |
| Are We Ready For An Agent-Native Memory System? | Zhou, Zhou, Han, Xu, Li, Li, Xiong, Wu | Jun 2026 | Systematic data-management-perspective evaluation decomposing agent memory into four modules — representation/storage, extraction, retrieval/routing, maintenance — and benchmarking 12 representative memory systems across 5 workloads / 11 datasets. Key finding: no single architecture dominates; effectiveness depends on whether the memory structure matches the workload bottleneck. Fine-grained ablations quantify effects on representation fidelity, retrieval precision, update correctness, long-horizon stability; localized maintenance beats global reorganization on cost. Code and paper-list released. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems | Bhandari, Wadhwani, Kumar, Narang | Aug 2026 | Takes one specific mechanism from western scrub jay episodic memory — per-memory, type-conditioned temporal decay — and operationalizes it as an auto-classified perishability coefficient in an external store, against the default that all memories are equally persistent. Each memory is a jointly-bound What–Where–When tuple carrying estimated perishability and a utility horizon, retrieved by query-adaptive scoring and revised retroactively at O(1) LLM calls per update. Introduces the Temporal Generalization Test (held-out retention intervals) and a Generalization Gap metric, on which ScrubJay-MEM is the only retrieval-based system with substantially positive GenGap (+0.108); on MemoryAgentBench EventQA-64k it improves F1 by +2.66 over Mem0 and +3.09 over Qwen3-Embedding-4B, and a decay ablation collapses GenGap by 5.7×. Unusually candid scoping: gains narrow under stronger backbones and reverse on fact-consolidation tasks 📄. ⚠️ Preprint. |
| Episodic Memory is the Missing Piece for Long-Term LLM Agents | Pink et al. (UT Austin) | Feb 2025 | Position paper arguing episodic memory — supporting single-shot learning of instance-specific contexts — is the critical missing capability. Defines 5 properties: temporal, instance-specific, single-shot, inspectable, compositional. |
| MAP: Modular Agentic Planner | Correa, Samwick, Gershman et al. (Microsoft Research, Harvard) | Oct 2023 / Nature Comms 2025 | Brain-inspired architecture decomposing planning into PFC-associated modules: conflict monitoring, state prediction, state evaluation, task decomposition, task coordination. arXiv:2310.00194. |
| SCL: Structured Cognitive Loop with Governance Layer | Kim | Nov 2025 | R-CCAM: Retrieval → Cognition → Control → Action → Memory. Soft Symbolic Control = governance layer applying symbolic constraints to probabilistic inference. Zero policy violations, complete decision traceability. |
| Procedural Memory Is Not All You Need | Wheeler, Jeunen | May 2025 | LLMs are constrained by reliance on procedural memory (pattern-driven tasks). For "wicked" learning environments with shifting rules and ambiguous feedback, different memory types are required. ACM UMAP '25. |
| Evaluating Theory of Mind in Multi-Agent LLM Systems | Kostka, Chudziak (Warsaw UT) | Sep 2025 | ToM and Internal Belief mechanisms are NOT universally beneficial — stronger models handle extra cognitive load well, but weaker models get confused. Model capability is the dominant factor. ICCCI 2025 ✅ (journal_ref). |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Demystifying Agent Skills: Why They Work—Until They Don't | Jiang, Huang, Xing, Wu, Gao et al. | Aug 2026 | Moves skill evaluation past aggregate success rates by isolating representation, outcome annotation, retrieval difficulty, and cross-framework robustness. Normalizes 8,135 trial records into a taxonomy of three categories and twelve skill-use modes. Core mechanism finding: skills work because noisy trajectories become procedural anchors that stabilize execution — 65.7% of cases, versus 4.5% for explicit knowledge injection (+6.06 points over Workflow Memory in matched comparisons); skills stabilize action, they do not inject missing facts. The scaling warning is the more useful half: retrieval is a separate bottleneck — as pools grow from 5 to 100, actual-use precision falls 29.6% → 3.3%. Exact ground-truth invocation is found to be neither sufficient nor necessary. Directly relevant to anyone scaling a skill/procedure catalog 📄. ⚠️ Preprint. |
| Muscle Memory for Agents: Compile not Merely Retrieve | Omran, Lanka, Zhang, Dixit | Aug 2026 | Position paper against the field's default pattern (store experience → retrieve at inference → let a general-purpose orchestrator interpret it), arguing it is the wrong default for personalization. Proposes compiling recurring user intent into purpose-built specialist agents as a memory paradigm distinct from retrieval, targeting the multi-turn tax where users repeatedly correct format, depth, and scope. Reference implementation is a four-phase pipeline (Harvest → Analyze → Augment → Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable specialists with two-stage trigger matching. On 90 held-out scenarios across five personas it wins 32 of 36 cases where a specialist fires (88.9%), at +2.05 personalization gain and a −0.28 accuracy cost on a 1–4 scale — note the win rate is conditional on a specialist firing 📄. ⚠️ Preprint. |
| SkillRouter: Skill-Based Routing for LLM Agents | Alibaba | Sep 2025 | Critical finding: skill BODY is the decisive routing signal (91.7% attention weight), NOT name (7.3%) or description (1.0%). Removing body causes 29-44pp degradation. BM25 on metadata alone scores 0%. Two-stage retrieve-and-rerank (1.2B params). |
| EvoSkill: Self-Evolving Skill Discovery | Sentient, Virginia Tech | Sep 2025 | Automates skill discovery via iterative failure analysis without model fine-tuning. Three agents: Executor, Proposer, Skill-Builder. Skill-merge outperforms single runs. Code |
| Trajectory-Informed Memory Generation for Self-Improving Agent Systems | Fang, Isahagian, Jayaram, Kumar, Muthusamy, Oum, Thomas (IBM Research) | Mar 2026 | Extracts typed actionable tips from execution trajectories: strategy tips (clean successes), recovery tips (failure-then-recovery), optimization tips (inefficient-but-successful). Closes the gap between episodic logs and reusable procedural knowledge — useful as a complement to EvoSkill and SkillRouter. |
| DeltaMem: Incremental Experience Memory for LLM Agents via Residual Trees | Tan, Zhang, Cao, Li, Chen (RUC) | Jun 2026 | Stores experience as residual deltas across two trees — goal-conditioned task skills and scene-level environment knowledge — so related episodes share a common root instead of duplicating content. Retrieval: failure-penalized similarity scan + root-to-match chain reconstruction; autonomous consolidation distills high-frequency paths into new roots. Cuts redundancy and retrieval conflicts; beats baselines across interactive environments 📄. Code released. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems | Xiong, Zhou, Wang, Gu, Tang, Li et al. (9 authors) | Aug 2026 | Targets a path defenses structurally miss: a poisoned memory is written once, retrieved repeatedly, promoted into shared memory, and reused by other agents — so a single write steers many later decisions and contaminates agents that never saw the original attack, while no malicious message ever crosses a visible communication edge at the moment of harm. Because existing safeguards inspect prompts, actions, or communication edges, they miss content that looks benign at write time but becomes harmful after retrieval. MAPLE-Guard monitors the memory lifecycle and places gates at write, retrieval, promotion, and cross-agent reuse, quarantining risky memories and blocking poisoned private memories before they enter shared memory. ASR 38.2% → 0.9% (LongMemEval) and 34.7% → 0.2% (AppWorld); multi-agent defense success rate 54.0%→74.3% and 42.5%→99.8% 📄. Code ⚠️ Preprint. |
| Governed Shared Memory for Multi-Agent LLM Systems (MemClaw / ArgusFleet) | Margalit, Cohen-Inger, Avram, Taig, Margalit | Jun 2026 | Formalizes the fleet-memory problem: unauthorized leakage, stale propagation, contradiction persistence, provenance collapse. Defines systems-level primitives — scoped retrieval, temporal supersession, provenance tracking, policy-governed propagation — implemented in MemClaw, a production multi-tenant memory service, evaluated live via ArgusFleet rather than a baseline comparison. Results: 100% correct depth-4 provenance reconstruction at sub-second/hop; zero cross-fleet leakage. Also reports real production bugs found via live eval: sub-tenant scope bypass on direct GET-by-id (disclosed/remediated) and a pipeline-ordering conflict where a synchronous near-duplicate gate can reject contradictory writes before the async contradiction detector runs. Conclusion: long-context retrieval alone is insufficient for production multi-agent memory 📄. |
| CoMAM: Collaborative Memory for Multi-Agent Systems | — | Sep 2025 | Collaborative RL framework modeling memory agents as sequential MDP with inter-agent dependencies. Group-level ranking consistency for coordinated memory operations. |
| DCM-Agent: Dual-Cluster Memory for Multi-Paradigm Ambiguity | Zhang, Wan, Zhang, Yang, Zhang, Wei, Liu | Apr 2026 | Tackles structural ambiguity in optimization problems where one problem admits multiple conflicting modeling paradigms. Training-free dual-cluster memory keeps competing paradigm solutions separated so the agent can pick or blend at inference time. Relevant to multi-agent settings where agents disagree on framing. |
April 2026 marks memory security emerging as a first-class research concern. As agent memories grow rich with user data, they become attack surfaces.
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| SkillJack: Persistent Skill Backdoors in Self-Evolving Agents | Ying, Wu, Wu, Zheng, Cheng et al. | Aug 2026 | Prior poisoning work only bites when a poisoned record is retrieved; SkillJack attacks the experience-to-skill pipeline instead, hijacking the agent's own learning process so poisoned experiences become durable behavioral artifacts. Three named properties: sanitization whitewashing (malicious intent obscured during skill extraction), cross-layer promotion (transient experience becomes persistent capability), and persistence isolation — 80.0% of skill-mediated attacks survive deletion of the original poisoned records, so source-record cleanup is not a remedy. On SkillX, safety detection drops from 98.5% on poisoned trajectories to 11.4% on the extracted skills, with attack success 56.2% (SkillX) / 89.2% (Anything2Skill) over 150 shared trajectories; some implanted skills unintentionally activate on benign queries. Motivates provenance-aware skill lifecycle protection. Code released (Tencent/AI-Infra-Guard) 📄. ⚠️ Preprint. |
| Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall (PPMF) | Xu, Xiao, Shao, Liu, Li | Jul 2026 | The attack-side twin of authority collapse. Identifies memory provenance laundering: during LLM-based consolidation, an external untrusted observation is rewritten as apparent user history or workflow support — preserving the action trigger while erasing the low-trust source that should limit its authority. Prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy consolidation. PPMF is lightweight middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. Vulnerable consolidated memories reach up to 1.000 ASR; with provenance, confirmation, and risk labels intact, no evaluated unauthorized high-risk action passes the gate while benign actions remain executable 📄. ⚠️ Preprint — arXiv comment reads "EMNLP2026 submitted" (submitted, not accepted). |
| MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair | Chen, Xie, Fu, Zhou, Yu, Xuan | Jul 2026 | Traces the same malicious semantics across persistence, downstream consequence, and selective repair — the repair axis prior poisoning benchmarks omit. 310 cases drawn from 48 realistic contexts (code/science, daily life, office work), each following a controlled Write–Execute–Forget protocol in an isolated runtime, with evidence-based adjudication across seven lifecycle checkpoints over a 24-configuration matrix (2 agent harnesses × 4 memory backends × 3 LLM backends). Across all 24: malicious memory persists in 84.2% of cases and the full Write–Execute chain succeeds in 50.3%; among successfully poisoned cases 59.6% complete the full Execute chain and 56.1% achieve selective repair 📄. ⚠️ Preprint. |
| Securing LLM-Agent Long-Term Memory Against Poisoning (TMA-NM) | Louck | Jun 2026 | Proves content- and lineage-based memory-poisoning defenses are malleable — an attacker can launder untrusted origin through the agent's own summarization, a trusted-tool echo, or manufactured corroboration, making poisoned content look benign and flipping its derivation edge to "trusted." Formalizes the malleability problem for the write-retrieve-act pipeline and proves a machine-checked separation theorem: write-time origin binding is necessary, non-malleable origin-bound authority with Sybil-resistant corroboration-gated elevation is sufficient. TMA-NM hits 0% attack success (vs. up to 68% laundering ASR for existing defenses) across 8 frontier models at full legitimate utility. Releases benchmark, harness, and machine-checked TLA+ models 📄. |
| When Agents Remember Too Much: Memory Poisoning Attacks on LLM Agents (GhostWriter / AM-Sentry) | Torres, Shrestha, Misra | Jul 2026 | Targets personal assistant agents — the convergence of conversational + action-planning agents handling sensitive data via untrusted sources. GhostWriter attacks in two phases (injection → activation) and achieves ~98% injection rate, ~60% activation rate against SOTA agents, exploiting the lack of security-focused memory governance. Proposes AM-Sentry (memory-saving policy + memory-retrieval screen) which dramatically cuts GhostWriter's success while preserving utility. |
| ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying | Lyu, He, Wang, Hu, Li, Chen, Li, Chen | Apr 2026 | First systematic privacy attack on agent memory. Combines data-distribution estimation of a victim agent's memory with entropy-guided adaptive querying to maximize leakage. Achieves up to 100% Attack Success Rate on extracting stored memories 📄 — substantially outperforming prior attacks. Establishes the need for privacy-preserving memory designs as a first-class research direction. |
| SSGM: Stability and Safety Governed Memory for LLM Agents | Lam, Li, Zhang, Zhao | Mar 2026 (rev May 2026) | Conceptual governance framework that decouples memory evolution from execution by enforcing consistency verification, temporal decay modeling, and dynamic access control before any memory consolidation. Provides a taxonomy of memory corruption risks: topology-induced knowledge leakage, semantic drift via iterative summarization, and consolidation hazards. Complementary defensive framing to ADAM. |
Wave 7 (June–July 2026) surfaces a distinct concern from privacy/extraction attacks: can you trust what memory itself contains and does to reasoning — corrupted consolidation, sycophantic over-reliance on stored user views, and stale/current facts silently coexisting in the same bank?
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| When Memory Becomes Authority: Benchmarking Authority Collapse at the Consolidation Boundary (AuthMem-Bench) | Zhan, Zhang, Guo, Zhao, Liu | Aug 2026 | Consolidation imposes an implicit authorization boundary — it decides whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. Names authority collapse: consolidation preserves the claim while erasing the source constraints governing its authorized use, so the memory implies greater authority than its source permits. AuthMem-Bench is a controlled paired benchmark holding the focal claim and downstream task fixed while varying only source authority. Across 7 consolidators × 7 backbones, collapse appears in 48 of 49 configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%; end-to-end, automatically predicted and persisted authority labels cut the observed rate from 16.9% to 0.0% with benign task success essentially unchanged 📄. Memory must preserve not only what was learned but the authority under which it may be reused. ⚠️ Preprint. |
| Beyond Similarity: Trustworthy Memory Search for Personal AI Agents (MemGate) | Zhang, Chen, Ma, Hu, He, Zhang, Liu, Yang, Zhang, Jia | Jun 2026 | Names memory search itself as a trust boundary: a semantically-similar memory can still be contextually inappropriate, causing cross-domain leakage, sycophancy, tool-call drift, or memory-induced jailbreaks. Evaluates A-Mem, Mem0, MemOS, and OpenClaw (real-world persistent-state agent) and finds long-term memory behaves as a durable control channel, not just a utility layer. Proposes MemGate — a 9M-parameter, 35.1MB plug-in inserted between the vector store and backbone LLM that applies a query-conditioned neural gate, turning raw similarity search into task-conditioned memory admission, with no LLM modification or inference-time judge required. Reduces memory-induced threats across frameworks/backbones while preserving utility 📄. |
| TRUSTMEM: Learning Trustworthy Memory Consolidation for LLM Agents with Long-Term Memory | Yang, Paul, Srinivasan, Kulkarni, Chappidi | Jun 2026 | Targets errors introduced by the write/revise/delete pipeline itself — omission, corruption, hallucinated content — which become persistent system-state failures once stored. A Memory Transition Verifier scores update transitions on coverage, preservation, faithfulness; preference pairs among candidate updates drive preference-guided RL to directly optimize consolidation behavior. SOTA on MemoryAgentBench, HaluMem, Mem-alpha; +12.14 F1 on HaluMem extraction; cuts omission/corruption/hallucination by 40.1% / 79.1% / 50.0% vs the strongest baseline per error type. |
| MemSyco-Bench: Benchmarking Sycophancy in Agent Memory | Xiang, Chen, Tang, Wei, Ning, Lin, Zhang, Su | Jul 2026 | Identifies memory-induced sycophancy: agents over-align with a user's stored prior views at the cost of factual accuracy or objective reasoning. Existing memory benchmarks check whether memories are stored/retrieved/updated correctly but not how they bias downstream reasoning. Five tasks probe whether an agent can reject memory as evidence, respect its applicable scope, resolve memory-vs-objective-evidence conflicts, track updates, and still personalize appropriately. First benchmark to treat memory reliance itself as a failure mode. |
| A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory | Shi, Tang, Tung | Jul 2026 | Names ghost memory: old, current, and transition facts coexist unmarked in the memory bank and get mixed during retrieval, misleading the answer model about what's true now. Argues memory should be evaluated at three decoupled levels — bank maintenance, retrieval, answer-time resolution — since final QA accuracy hides where the failure actually occurs. ATMA is a state-aware overlay that keeps superseded/transition records, builds evidence packets for the query's requested state view, and exposes current/historical/transition labels to QA. On LTP (new LoCoMo Temporal Plus benchmark), Graphiti+ATMA improves conflict accuracy +0.240 absolute; on LoCoMo, raises temporal F1 from 0.0295 to 0.1705. |
Wave 9 (July–August 2026) surfaces the window's strongest convergent signal: in roughly four weeks, multiple independent groups imported database transaction semantics wholesale into agent memory. The shared premise — a memory write is not a belief commit — is forced by the fact that stored memory now drives irreversible external actions. Staged writes, snapshot isolation, versioned heads, natural-language rollback, and cascading repair of derived records are becoming memory-layer primitives rather than database trivia.
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| MemTX: Transactional Belief Commit for Stateful Agent Memory | Li, Wang, Lu, Chen, Li, Song, Zheng, Cai | Jul 2026 | The cluster's strongest entry and the source of its thesis. In shared-memory multi-agent settings one agent's write becomes another's premise, and eventually a tool call with real side effects — yet current systems treat every accepted write as immediately actionable truth, so a polluted tool result or a teammate's half-finished note can silently drive an irreversible action. Each record carries evidence, permissions, provenance, and validity; writes are staged inside snapshot-isolated transactions admitted by a validate-and-commit pipeline; irreversible tool calls are gated on in-flight belief state; and retracting a belief triggers typed cascading repair of its derived records and tool side effects. Two invariants — action-safety gating and cascade-repair completeness — are machine-checked via property-based testing and bounded exhaustive enumeration of 5.5M protocol states, with zero violations. Across five backbones from three model families it leads all eight baselines with paired-McNemar significance on four and ties the strongest fifth, and is the only method with zero downstream harm on every backbone 📄. "Backbone capability does not substitute for commit discipline." ⚠️ Preprint. |
| MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery | Cui, Tang, Yao, Meng, Ma, Jia | Jul 2026 | A governance layer outside the answer model, supplying the transaction boundary existing systems lack: Ordered PatchTest verifies whether an update is actually supported by its source, a Temporal Resolver selects the visible version when facts conflict, and a durable snapshot journal restores application-visible state after a fault. On an item-disjoint audit it accepts all 60 supported originals and rejects all 179 hard negatives; under persistent multi-key faults on LongMemEval-S and LoCoMo states it restores the complete declared active map without knowing the actual physical write set. Highest average F1 across all twelve answer-model configurations on MemoryAgentBench FactConsolidation, beating Dense by 17.06–24.07 points in five representative settings 📄. ⚠️ Preprint. |
| ChronoMem: Version Control and Semantic Rollback for LLM Agent Memory | Su, Xu, Zuo, Bertino | Jul 2026 | Attacks forward-only evolution: memory systems continuously accumulate, consolidate, and overwrite with no principled mechanism to inspect, version, or revert — leaving agents brittle under corrections, concept drift, and corruption, particularly after they have already been exposed to subsequent information. ChronoMem commits whole-memory snapshots at every write, maintains structured version histories, and maps natural-language undo intents to concrete historical versions through hybrid lexical + semantic retrieval, rank fusion, and reranking. Introduces a post-exposure counterfactual protocol: can the agent answer queries and summarize history as if future updates had never occurred? Integrated into Google's production-ready, open-source Agent Development Kit; claims the first open-source system and benchmark for systematic semantic global memory rollback. ⚠️ Preprint. |
| TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management | Xiao, Xu, Zhang, Chen, Shi | Aug 2026 | Retires the binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification — choices that may share a label while producing fundamentally different memory states. TARL maps each statement to one of five executable actions, identifying the affected memory, resolving its temporal scope, comparing source reliability, and updating accepted, pending, and rejected ledgers. Trained by comparing the memory states produced by alternative update operations. Ships TARL-Mem, a benchmark with fine-grained action labels and next-state targets; improves action prediction and state recovery, reduces memory pollution, preserves conflicting evidence, and limits cumulative corruption 📄. ⚠️ Preprint. |
| Paper | Authors | Date | Key Contribution |
|---|---|---|---|
| Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems | Pollertlam, Kornsuwannawit | Aug 2026 | The cost-side counterpart to accuracy benchmarking: Mem0, Hindsight, and Mastra Observational Memory against two reference strategies (fixed-size rolling window, full-transcript resubmission), two backbones, conversations up to 400 turns, every cost measurement paired with accuracy on 665 LoCoMo questions. Three findings, all awkward for the field: (1) serving cost cannot be predicted from conversation length and message size alone — a regression that tracks the reference strategies closely misses the memory systems by 18–69%, cost being driven by internal memory behavior; (2) break-even against full transcript is highly sensitive to system and backbone, ranging from the first tens of turns for the cheapest to never within 400 turns for the most expensive; (3) no system wins on both axes — accuracy spans 21–54%, and backbone choice drives cost as much as the memory system does 📄. ⚠️ Preprint. |
| Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents | Alake, Bernardis, Cayet, Engel, Hilloulin, Hong, Hosler, Kavantzas, Kossyk, Le, Patra, Talamadupula, Venzin (Oracle) | Jul 2026 | Database-native memory substrate built on Oracle Database, framed as a full lifecycle: ingestion → extraction → consolidation → retrieval → summarization → revision/removal, with a layered active-core / passive-store split and explicit scope control across users, agents, threads. Reports 93.8% LongMemEval accuracy using ~10.7x fewer tokens than flat-history baselines. Notable as a large enterprise-vendor entry (13 authors, Oracle) validating the memory-as-lifecycle framing that most of this list converges on independently 📄. |
| Beyond the Context Window: Cost-Performance Analysis | Pollertlam, Kornsuwannawit | Sep 2025 | Compares Mem0-style fact-based memory vs long-context LLMs on LongMemEval, LoCoMo, PersonaMemv2. Break-even: memory system becomes cheaper after ~10 interaction turns at 100K context. Long-context wins on factual recall but memory is competitive on reasoning. |
| Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads | Omri, Gan, Broveak, Geens, He, Pentland, Verhelst, Weissman, Tambe (Stanford / MIT / KU Leuven) | Jun 2026 | The first systems-level characterization of agent memory: (1) a system-oriented taxonomy along four axes; (2) a phase-aware profiling harness attributing cost to construction, retrieval, and generation; (3) characterizes 10 representative systems across two benchmark suites, showing how design choices shift cost across write vs read paths; (4) derives 10 system recommendations (construction scheduling, capability floors, amortization via query volume, freshness-latency tradeoffs, fleet-scale management). The production "memory economics" lens. 📄 ⚠️ Preprint. |
Most agent memory research ignores 50 years of neuroscience. This section bridges that gap — presenting the biological mechanisms that may underlie what LLM memory papers have independently re-discovered (forgetting, gating, dual-phase encoding), alongside the spiking neural network implementations that model them directly. Cross-references to LLM papers below are 🔬 editor's synthesis.
| Paper / Project | Authors | Date | Key Contribution |
|---|---|---|---|
| A Bio-realistic Synthetic Hippocampus for Robotic Cognition | Talanov et al. | Oct 2025 | Synthetic hippocampal architecture with dual-phase operation: online sensorimotor encoding during wake, offline consolidation via SWR-triggered replay during sleep. SNN on neuromorphic substrates (≤1W). Goal-prioritised plasticity prevents catastrophic forgetting. Models the biological mechanisms that parallel what SleepGate and CraniMem implement in software 🔬. BioNanoScience. Open access. |
| The Memristive Implementation of the Hippocampus: A Hypothesis | Talanov et al. | Aug 2025 | Hardware-level hippocampal memory using stochastic polycrystalline nano-fiber mesh as memristive substrate. Key insight: the inherent randomness of material structure mimics probabilistic biological synaptic networks. Demonstrates resistive switching tunability for implementing dynamic memory functions — bidirectional replay, synaptic up/down-scaling, consolidation. BioNanoScience. Open access. |
| Simulation of Serotonin Mechanisms in NEUCOGAR Cognitive Architecture | Talanov, Gafarov, Vallverdú et al. | 2018 | Maps neuromodulatory mechanisms to computational models: dopamine → attention, serotonin → inhibition. The "cube of emotions" model. Demonstrates that mammalian emotional-state control via monoamine neurotransmitters can be re-implemented computationally. Foundation for neuromodulator-gated memory admission. Procedia Computer Science. |
| tinyHippo | Talanov | Active | CA1 + CA3 hippocampal microcircuit simulation in NEST with Izhikevich neurons. Implements bidirectional replay, theta-modulated encoding/retrieval phase separation, and SWR-triggered consolidation. The biological reference implementation — validates the mechanisms that Membrain engineers and SleepGate abstracts. MIT license. |
| Membrain | Fatykhov | Active | Neuromorphic memory bridge using FlyHash encoding and BiCameralMemory SNN (Nengo/Voja learning). Hopfield-style attractor dynamics for pattern completion (tested up to 20% noise in PoC). Stochastic consolidation with SleepSignal. gRPC API for agent integration. The engineering abstraction layer between biological models (tinyHippo) and cognitive agents (Nous). |
The Stack: These projects form a natural hierarchy — tinyHippo (biological model, validates mechanisms) → Membrain (engineering abstraction, SNN service) → Nous (cognitive agent, consumes memory). The papers provide the theoretical foundation; the repos provide working implementations.
Why this matters for agent memory 🔬: The LLM papers in this list independently converge on mechanisms that neuroscience has studied for decades. The following correspondences are editor's synthesis — the LLM papers don't explicitly cite these neuroscience sources, but the structural parallels are striking:
Papers describe mechanisms; these repos run them. Inclusion bar: (1) implements a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance; (2) actively maintained and (3) an OSI-approved license — both enforced for the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where a reader may actually vendor the code; author code for papers and benchmark/dataset repos is held to a different standard — reproduction value — so it may carry restrictive or absent license terms and may go quiet after publication, with both facts flagged inline ⚠️; (4) and either backs a paper in this list, has real production adoption, or fills a mechanism niche nothing else covers. Agent harnesses that merely bundle a vector store, general-purpose vector databases, and repos whose only claim is a self-reported benchmark ranking are deliberately excluded. Stars, license, and last-push verified via the GitHub API on 23 Aug 2026 — a snapshot for triage, not a ranking.
| Project | Stars | License | What it actually implements |
|---|---|---|---|
| mem0 | 63.8k | Apache-2.0 | Extraction-based fact memory: an LLM pass distills durable facts from raw dialogue, then reconciles them against the existing store via add / update / delete / noop. The most widely deployed reference for the extract-then-reconcile write path. Paper: Mem0. |
| Graphiti (Zep) | 30.2k | Apache-2.0 | Bi-temporal knowledge graph: edges carry validity intervals and are invalidated, not overwritten, so superseded facts remain queryable with time bounds. The closest production analogue to Memory Transactions, Versioning & Recovery. |
| cognee | 30.2k | Apache-2.0 | ECL (extract → cognify → load) pipeline turning documents and conversations into a typed knowledge graph plus vector index, with explicit prune/forget operations over the graph rather than append-only growth. |
| Letta | 24.4k | Apache-2.0 | The MemGPT lineage: OS-style paged memory where the agent edits its own core memory block through tool calls and pages the remainder to archival storage. Self-editing memory as the mechanism. Paper: MemGPT. |
| Hindsight | 20.9k | MIT | Consolidation as an explicit policy layer over four levers — importance, merge, decay, eviction — instead of an unbounded append log. Industry framing already cited in Adjacent Research; paper: arXiv:2512.12818. |
| MemOS | 10.9k | Apache-2.0 | "Memory OS": a MemCube abstraction unifying plaintext, activation (KV), and parameter memory, with scheduling and cross-task reuse between them. Rare in treating KV-cache and long-term store as one substrate. |
| Honcho | 6.8k | AGPL-3.0 | Memory as a representation of a person, not a transcript: background reasoning maintains per-peer models, queried through a dialectic API rather than similarity search. |
| MemMachine | 3.2k | Apache-2.0 | Clean two-tier service — episodic memory plus a durable user profile — with pluggable stores and a server/client split. A readable reference implementation if you are building your own. |
| MemoryOS | 1.6k | Apache-2.0 | EMNLP 2025 Oral ✅. Short / mid / long-term stores with heat-based promotion and eviction between tiers — one of the few systems whose admission and consolidation policy is inspectable code rather than a prompt. |
| redis/agent-memory-server | 308 | Apache-2.0 | Explicit working-memory (token-bounded, session-scoped) vs long-term memory split, with automatic extraction at the boundary. MCP + REST, Redis-native. |
Reference implementations for papers already indexed in this list — verified to exist and to correspond to the cited work.
| Paper (in this list) | Code | Notes |
|---|---|---|
| xMemory | HU-xiaobai/xMemory | ⭐120 · MIT. Author release for Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation. |
| AgeMem | y1y5/AgeMem | ⭐37 · ACL 2026 SAC Highlight ✅. Unified long/short-term memory management learned as tool operations. |
| CraniMem | PearlMody05/CraniMEM | ⭐7 · gated-and-bounded memory implementation; project page. ⚠️ No license file — all rights reserved by default; clear reuse with the authors. |
| HeLa-Mem | ReinerBRO/HeLa-Mem | ⭐16 · ACL 2026 Main ✅. Hebbian associative distillation, reproduces the LongMemEval-S numbers. ⚠️ No license file — all rights reserved by default. |
| StructMemEval | yandex-research/StructMemEval | ⭐11 · Apache-2.0. Raw benchmark data + supplementary code from Yandex Research; authors label it work in progress. |
| memorywire | mthamil107/memorywire | ⭐2 · Apache-2.0. Reference implementation of the wire format, including the provenance-based purge/quarantine path for poisoned stores. |
| A-MEM | agiresearch/A-mem | ⭐1.2k · MIT. Zettelkasten-style memory that links and evolves notes on write. Paper reproduction lives in WujiangXu/A-mem (NeurIPS 2025 ✅). Last push Dec 2025 ⚠️. |
Single-binary or file-backed systems, usually MCP servers, where the store is inspectable by a human and runs without a cloud dependency.
| Project | Stars | License | What it actually implements |
|---|---|---|---|
| agentmemory | 27.3k | Apache-2.0 | Full lifecycle memory for coding agents: confidence scoring, decay and consolidation, knowledge graph, hybrid search, ~20 agent adapters. ⚠️ Headline ranking claims are self-reported. |
| Engram | 6.1k | MIT | Agent-agnostic Go binary — SQLite + FTS5, MCP/HTTP/CLI/TUI, explicit conflict handling. Deterministic lexical recall as a deliberate counter-position to embedding everything (cf. ReFind in Retrieval & Recall). |
| memsearch (Zilliz) | 2.5k | MIT | Markdown files as the source of truth, Milvus for hybrid retrieval, typed episodic/procedural memory. Human-readable store = auditable memory. |
| Nocturne Memory | 1.3k | MIT | Rollbackable, visually inspectable graph memory server for MCP agents — edits are reversible operations, not blind writes. The practical counterpart to MemTX/ChronoMem-style recovery. |
| deja-vu | 1.0k | MIT | No write path: indexes the session transcripts 34 coding agents already write locally (lexical by default; optional deja embed may send text to an external embedding API) and serves them to any of the others — including pre-install history, so no cold start. Read-side mechanisms: outcome-aware ranking (a session that ended in a fix outranks one that gave up), redaction at ingest, per-project trust policy, forget/unforget. LoCoMo/LongMemEval drivers in-repo. |
| Vestige | 608 | AGPL-3.0 | Local-first Rust MCP server combining FSRS-6 decay, spreading activation, contradiction inspection, active suppression, Receipt Lock, and an inspectable dashboard. Built for causal recall — the quiet change that broke today, not the lookalike. |
| Mnemon | 507 | Apache-2.0 | Single Go binary, four-graph knowledge store with intent-aware recall, importance decay, and automatic deduplication. No API keys, one setup command. |
| Caura (ex-MemClaw) | 439 | Apache-2.0 | Governed shared memory for agent fleets: trust tiers, keystone policies, audit trails, multi-tenant scoping. The implementation counterpart to Memory Trust & Integrity and Multi-Agent Memory. |
| mengram | 189 | Apache-2.0 | Typed semantic / episodic / procedural memory where procedures are derived from recorded failures — one of the few OSS systems implementing the failure → procedure loop from Skill & Procedural Memory. |
| causal-memory | 66 | Apache-2.0 | Local-first Rust MCP server (also CLI + PyO3 bindings) where the store is one SQLite file holding facts, temporal state, and typed decision→outcome edges (caused / enabled / prevented). Spreading activation runs both ways — prevented edges carry negative activation, a GABA analogue — alongside Hebbian co-occurrence, Q-value dynamics, write-time gating, and immutable SWR-style consolidation. Reproducible benchmark harnesses (CausalEval, LoCoMo, LongMemEval) live in-repo; Hermes Agent provider plugin included. |
| Tree Ring Memory | 13 | MIT | Project-scoped Rust CLI lifecycle layer: SQLite/FTS explainable recall, deterministic consolidation, forgetting, redaction, audit, framework discovery. Memory that ages instead of accumulating raw transcript. |
| kgai | 1 | MIT | Append-only, content-addressed decision log for dev teams: every change is a new event that supersedes the prior one through an explicit link with a reason, rejected paths stay queryable (temporal versioning), recall returns only decisions in force. Deterministic graph projection, lexical recall with no embeddings, and per-writer shard sync over the team's own S3 bucket (no server). Repo-supplied config is gated behind an explicit trust approval (governance). Claude Code plugin plus a Go CLI. |
| Hyperconsciousness | 1 | MIT | Rust knowledge store with signed, encrypted append-only records and scoped, expiring grants for MCP access. Delegated grants cannot widen scope, add actions, or outlive their parent. Governance at the retrieval boundary; does not sandbox processes running as the owner. ⚠️ Developer alpha; no independent security audit. |
Neuromorphic reference implementations — tinyHippo and Membrain — live in Neuromorphic & Bio-Inspired Memory.
| Project | Stars | License | What it measures |
|---|---|---|---|
| LongMemEval | 1.0k | MIT | ICLR 2025 ✅. 500 questions across five long-term memory abilities over long interactive histories. The de-facto standard — and the number most vendor claims cite. |
| LongMemEval-V2 | 132 | Apache-2.0 | Official successor, released 2026, addressing saturation and contamination in the original. |
| LoCoMo | 1.1k | see repo ⚠️ | Very long-term conversation benchmark (ACL 2024 ✅). Widely used dataset; frozen ⚠️ — last push Aug 2024, cite the data, don't expect maintenance. |
| MemoryData | 139 | none ⚠️ | Unified harness — 4 benchmark families, 22 method presets, one runtime — built specifically to make heterogeneous memory systems comparable. Directly attacks the fragmentation problem Evaluation & Benchmarks documents. |
| GateMem | 197 | MIT | Memory governance under multiple principals sharing one store: utility, access control, and active forgetting measured together instead of accuracy alone. |
| HaluMem | 155 | CC-BY-NC-ND-4.0 ⚠️ | Operation-level hallucination benchmark: scores extraction, updating, and answering separately, exposing errors that end-to-end QA accuracy hides. License is declared by README badge only — no LICENSE file in the repo; NC-ND terms forbid commercial use and derivatives, so treat it as read-only unless you clear it with the authors. |
| Mem-Gallery | 103 | MIT | ACL 2026 Main ✅. Multimodal long-term conversational memory for MLLM agents. Last push Jan 2026 ⚠️. |
| STATE-Bench | 77 | MIT | Microsoft. Memory-agnostic enterprise workflows measuring whether an agent improves with experience, not whether it recalls a fact. Blog write-up in Adjacent Research. |
| Resource | Stars | Why it's here |
|---|---|---|
| Agent Memory Techniques | 926 | 30 runnable notebooks covering buffers, vector stores, KGs, episodic/semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, and LoCoMo evaluation. The best hands-on on-ramp before reading the papers above. |
| Awesome-AI-Memory (IAAR Shanghai) | 1.2k | Broad bilingual knowledge base spanning LLM memory and agent memory, including engineering frameworks and applications. |
| Awesome-Agent-Memory (TeleAI) | 596 | Systems + benchmarks + papers for LLM and MLLM memory — stronger multimodal coverage than this list. |
| Awesome-GraphMemory | 339 | Survey-grade depth on graph-based agent memory specifically. |
| Awesome-Agent-Memory-Papers | 229 | Paper-only index with a browsable website. |
| OpenDataBox/awesome-agent-memory | 54 | Same name, different list — a four-axis paper taxonomy, companion to the MemoryData harness above. |
These papers and projects aren't solely about agent memory but contribute relevant architectural ideas:
| Paper / Project | Date | Relevance |
|---|---|---|
| MemTools: A Unified Research Framework for Interoperable Agent Memory — Zhao, Chen, Liang, He, Wang, Zhao, Liu | Jul 2026 | Targets architectural fragmentation as the thing blocking systematic memory research: implementations couple lifecycle stages together, entangle evaluation logic with specific datasets, and barely support heterogeneous memory types. MemTools standardizes the memory lifecycle through declarative data contracts, enabling interchangeable assembly of components across systems, and orthogonally separates benchmark datasets from execution protocols so evaluation can be reconfigured independently of data. Adds a unified computational interface for coordinating symbolic, neural, and multimodal representations in one runtime. ⚠️ Authors label it work in progress; the evaluation demonstrates isolation of design variables rather than task-level performance. |
| MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories — Lovén, Sauvola, Riekki, Tarkoma | Aug 2026 | Agents share transport and can call each other's tools, but have no protocol for reconciling what they know. MELD admits every incoming claim through a five-outcome procedure (insert, merge, relate, conflict, reject) decided from three signals — scoped claim-key identity, embedding similarity, and a natural-language- inference verdict — under context and freshness gates, acting through exactly one auditable, authenticated Patch as the sole object that mutates state, with a per-claim status CRDT over standard pub/sub. The design stance is the notable part: MELD does not adjudicate truth — a detected contradiction is preserved for later adjudication, never silently resolved. Merge classifier separates at AUC 0.968 with a 0.013 false-merge rate; the status CRDT reconverges in 30/30 real partition-heal trials where last-writer-wins manages 11/30; recall-non-inferior to a centralized store at ~11% less live storage; ~3× fewer messages at matched recall. Evaluated on a real continuum spanning operator-grade 5G edge, national HPC, and a local tier 📄. ⚠️ Preprint; code and data on Zenodo. |
| memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations — Munirathinam | May 2026 (rev Jul 2026) | JSON-Schema 2020-12 wire format for 5 memory ops (remember/recall/forget/merge/expire) × 4 memory types over mem0/Letta/Cognee/Zep/pgvector backends, with an optional HITL governance channel. Not a new algorithm — a packaging of RRF/FSM/consolidation into a vendor-neutral protocol meant to compose with MCP. v3 adds a poison-recovery eval (PurgeBench) showing its provenance field is the strongest lever for recovering a poisoned store — relevant complement to TMA-NM/GhostWriter above. |
| EverMemOS — Liu, Bai, Chen et al. | Jan 2026 | Self-Organizing Memory Operating System for structured long-horizon reasoning. |
| AriGraph — Anokhin, Radionov et al. | IJCAI 2025 ✅ | Knowledge Graph World Models with episodic + semantic memory. Proceedings |
| Memory Matters More Than You Think — Zhou, Li et al. (Renmin U.) | Jan 2026 | Event-centric memory with Logic Map for agent searching and reasoning. |
| The Consolidation Problem in Agent Memory — Vectorize (Hindsight) | May 2026 | Industry framing of consolidation as a policy layer with four levers — importance, merge, decay, eviction — benchmarked across Mem0, Zep, Letta, LangChain, Hindsight. |
| Introducing STATE-Bench — Microsoft (repo) | May 2026 | Open-source, memory-agnostic benchmark measuring whether agents improve with experience on stateful enterprise tasks (travel, support, shopping) — not just recall. |
| State of AI Agent Memory 2026 — Mem0 | 2026 | Industry survey of benchmarks, architectures, and six open problems (temporal abstraction, cross-session structure/identity, app-level eval, privacy/consent, staleness). |
These papers, taken together, converge on several architectural principles:
Papers: A-MAC, CraniMem, SleepGate, ACC
The strongest systems filter before storing. CraniMem uses goal-conditioned cosine gating; A-MAC scores across 5 dimensions (utility, confidence, novelty, recency, type). The alternative — storing everything and retrieving selectively — leads to noise accumulation and proactive interference.
Papers: xMemory, SYNAPSE, AdaMem, CraniMem, Mem0, Episodic Memory
Every serious architecture separates memory by type (episodic, semantic, procedural, working). xMemory's 4-level hierarchy shows 44.91% of memories get retroactively reassigned — proving flat structures fundamentally fail. The human brain doesn't use one store; neither should your agent.
Papers: SleepGate, CraniMem, SYNAPSE, A-MEM
Raw episodic traces should consolidate into durable semantic knowledge over time. SleepGate's sleep cycles, CraniMem's replay-based promotion, and SYNAPSE's spreading activation all implement variants of this. Without consolidation, memory becomes a write-only log.
Papers: SleepGate, Procedural Memory, xMemory
SleepGate reduces proactive interference from O(n) to O(log n) through learned forgetting. This is the most under-implemented capability in agent frameworks — most systems only add memories, never remove or decay them. Forgetting is what keeps memory useful rather than just large.
Papers: xMemory, Memory Survey, Anatomy of Agentic Memory
The clearest thesis across this literature: retrieval-augmented generation is a retrieval strategy, not a memory system. Agent memory requires admission, evolution, consolidation, forgetting, typed storage, and context-aware retrieval. RAG addresses only the last item.
Papers: Anatomy of Agentic Memory, StructMemEval
Current benchmarks are saturated, metrics misalign with semantic utility, and results vary wildly by backbone model. StructMemEval reveals agents can't organize memory autonomously. The field needs better evaluation before it can claim progress.
Papers: AgeMem, EMPO², MemFactory, Memory-R1
The dominant shift in 2026: from hand-coded admission/retrieval rules to reinforcement-learned memory policies. AgeMem exposes all memory operations as tool-based actions and learns when to use them via GRPO. EMPO² takes this further by using memory for exploration, not just recall. MemFactory provides the infrastructure to train these policies. This is the most consequential paradigm shift since the field recognized that RAG ≠ memory.
Papers: ADAM, FSFM
April 2026 surfaced the first systematic memory-extraction attack (ADAM, up to 100% ASR) — and the first forgetting frameworks (FSFM) that explicitly frame selective forgetting as a privacy primitive, not just a retention/efficiency one. Expect memory threat-modeling, audit interfaces, and unlearning APIs to follow.
Papers: Experience Compression Spectrum, Externalization in LLM Agents
The memory and skill-discovery communities have been solving the same problem (extract reusable knowledge from interaction traces) without citing each other — cross-community citation rate is below 1%. The Compression Spectrum reframes both as points on a single axis (5–20× → 50–500× → 1000×+ compression), and identifies the missing diagonal: no current system supports adaptive cross-level compression.
Papers: MemReader (active extraction), SCG-MEM (schema-constrained generation), HeLa-Mem (Hebbian distillation)
Two paradigm shifts in April 2026: (1) MemReader moves extraction from passive transcription to an RL-learned WRITE/DEFER/RETRIEVE-CONTEXT/DISCARD policy. (2) SCG-MEM replaces dense retrieval with schema-constrained decoding (Piaget-style assimilation vs accommodation). Together they suggest the future of memory is generative-constructive, not retrieve-and-paste.
Papers: MRAgent, MemPro, DeltaMem
Wave 6's headline shift: stop treating memory access as a static retrieve-then-reason step. MRAgent folds LLM reasoning into graph traversal (Cue-Tag-Content + active reconstruction), iteratively pruning paths on accumulated evidence. MemPro generalizes the idea to the whole pipeline — the memory system itself becomes an evolvable program that rewrites its own construction/retrieval logic. DeltaMem reconstructs each experience on demand by composing a root-to-leaf residual chain rather than storing flat copies. Memory is increasingly constructed at access time, not fetched.
Papers: Agent Memory Characterization, Beyond the Context Window, Personalize-then-Store
For two years memory papers reported accuracy; few reported cost. Wave 6 changes that. The first systems characterization (Omri et al.) builds a phase-aware cost profiler and a 4-axis taxonomy, deriving 10 deployment recommendations across the write and read paths. Personalize-then-Store reframes admission as a budget problem — gating transient sessions to spend the memory budget where it pays off. The question is shifting from "does memory help?" to "what does memory cost, and where does it amortize?"
Papers: TRUSTMEM, MemSyco-Bench, A-TMA
Wave 6 established memory as an attack surface (ADAM). Wave 7 shows a second, orthogonal failure class: memory that's technically retrieved correctly but still misleads. TRUSTMEM targets self-inflicted corruption during consolidation (omission, hallucination). MemSyco-Bench targets over-reliance — agents favoring a user's stored prior view over current evidence. A-TMA targets ghost memory — stale and current facts silently coexisting in the same bank. None of these are extraction attacks; they're reliability failures baked into ordinary write/retrieve operations.
Papers: MemProbe, Are We Ready For An Agent-Native Memory System?
MemProbe's headline finding — task completion nearly saturates even for a memoryless baseline, while structured state recovery from the memory artifact itself stays around 0.6 — is a direct rebuke of end-to-end accuracy as a memory metric. "Are We Ready" pushes the same point systemically: decomposing 12 memory systems into 4 modules shows no architecture dominates, and effectiveness is bottleneck-specific, not a single leaderboard number. Both extend Anatomy of Agentic Memory's Feb-2026 critique from theory into large-scale measurement.
Papers: What to Keep What to Forget (Rate–Distortion View), Is Agent Memory a Database?
Two Wave 7 papers step back from architecture-of-the-week and ask for formal foundations. The rate–distortion paper shows KV-cache eviction, prompt pruning, bounded recurrent state, and agent memory consolidation are the same budget-constrained retain/discard decision, and that all four fail identically (irreversible, pre-query, recency/attention-biased). "Is Agent Memory a Database?" argues memory correctness is a property of the state trajectory, not individual records — no record-level storage system can satisfy it, motivating the GEM formalization. Both suggest the field's next gains come from theory, not another heuristic gate.
Papers: AuthMem-Bench, PPMF
Two independent papers landed the same finding within days, from opposite directions. Summarization is lossy in a way nobody was measuring: it preserves the claim and drops the permission. AuthMem-Bench measures it as an accident — authority collapse in 48 of 49 configurations, a 50.3% mean unauthorized-action rate for collapsed memories. PPMF weaponizes it — laundering an untrusted observation into apparent user history, preserving the trigger while erasing the low-trust source, up to 1.000 ASR. Both fixes are the same shape: persist authority as first-class metadata and match action risk against it at tool-call time. This is distinct from insight #13, which asks whether a memory is true; this asks what a memory is allowed to authorize.
Papers: MemTX, MemTxn, ChronoMem, TARL
The transactions section in one line. Agent memory is acquiring database semantics — staged writes, snapshot isolation, versioned heads, rollback, and cascading repair of derived records and tool side effects. The forcing function is that memory now drives irreversible external actions, so the cost of an unsound write is no longer a bad answer but a real-world consequence. Note the direction of travel: TARL replaces a binary Write/Hold with five actions; ChronoMem adds an undo; MemTX gates tool calls on in-flight belief state. All three treat the moment of writing as provisional by default.
Papers: ReFind, Filesystem-Based Memory, Harness the Memory, Keep It InMind
The sharpest contrarian thread this list has carried. ReFind beats HippoRAG 2 (58.2 vs 53.2) with no semantic structure at all, on a matched backbone. Filesystem-Based Memory finds organization buys retrieval economy, not accuracy — and erodes as stores grow. Harness the Memory finds no substrate dominates and that excessive retrieval actively harms sequential decision-making. Against all three, Keep It InMind argues the framing itself is wrong: the failure isn't structure versus no-structure but the query-conditioned retrieval interface — 84.0% with memory in context collapsing to ≤14.4% when the same memory must be retrieved. Read together they pose an uncomfortable question for most of this list: is the field over-investing in write-time structure to compensate for a broken read-time interface?
This list grows as the field grows. To contribute:
Papers should be peer-reviewed, at reputable workshops (e.g., ICLR MemAgents), or highly cited preprints. We prioritize papers with novel architectural ideas over incremental benchmark improvements.
For repositories (the Implementations section), the bar is different: a repo must implement a memory mechanism beyond vector-store RAG — admission, consolidation, forgetting/decay, typed multi-store, temporal versioning, or governance — and either back a paper in this list, show real adoption, or fill a niche nothing else covers. Star counts are context, not a criterion. Self-reported benchmark rankings are not evidence; a reproducible harness is. Active maintenance and an OSI-approved license are required only in the two reusable-implementation subsections (Memory Layers & Frameworks, Local-First & Coding-Agent Memory), where the code is meant to be vendored. Author code for papers and benchmark/dataset repos is judged on reproduction value instead: a license must still be disclosed, and where it is absent or non-OSI (e.g. NC-ND, or no LICENSE file at all) — or where the repo has been quiet for more than six months — the row is flagged ⚠️ rather than dropped. A frozen benchmark that everyone cites still earns its place; it just gets its last-push date stated.
This list is released under CC0. The papers themselves retain their original licenses.