Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
Python
56
1 commits
updated Jul 14, 2026
As a trajectory grows, the state that should drive the next action — requirements, environment facts, prior attempts, diagnoses, open subgoals — gets buried in the context window or pushed past it, and stops influencing decisions when it's needed. We call this behavioral state decay, and treat memory as an active intervention rather than passive retrieval. Proactive Memory Agent runs a plug-and-play memory agent alongside an unmodified action agent: it reads the recent trajectory, manages a structured memory bank, and decides whether to inject a memory-grounded reminder to the action agent or stay silent.
Rendered agent trajectories (open in browser via raw.githack.com):
| Task | Baseline | Memory (action agent trajectory) | Memory (memory agent trajectory) |
|---|---|---|---|
| git-multibranch | action agent trajectory | action agent trajectory | memory agent trajectory |
| regex-log | action agent trajectory | action agent trajectory | memory agent trajectory |
| sqlite-with-gcov | action agent trajectory | action agent trajectory | memory agent trajectory |
| hf-model-inference | action agent trajectory | action agent trajectory | memory agent trajectory |
| adaptive-rejection-sampler | action agent trajectory | action agent trajectory | memory agent trajectory |
System integration. The action agent (left) interacts with the environment; the memory agent (right) runs alongside, observing a sliding window of recent steps and the current memory store. At every N steps the memory agent is invoked to update the bank and optionally inject a context reminder into the next action-agent call.

Memory agent internals. Phase 1 manages the memory bank (status, knowledge,
procedural entries) through tool calls. Phase 2 reads the updated bank and either
emits a <context_for_action> reminder or <no_intervention/>.

git clone https://github.com/yifannnwu/proactive-memory-agent.git
cd proactive-memory-agent
# 1. Install the vendored Harbor (required by the Terminal-Bench runner)
pip install -e external/harbor
# 2. Install this package
pip install -e .
Python 3.12+. Uses litellm for LLM calls (works with OpenRouter, OpenAI,
Anthropic, local vLLM, etc.).
Quick import check:
python -c "import memory_agent; print('OK')"
python -m memory_agent.cli --help
The Terminal-Bench runner uses Harbor + Enroot for sandboxed shell execution
(chosen because our cluster is rootless and shares an FSx cache of pre-built
.sqsh images). Both configs run the full terminal-bench@2.0 set (89 tasks).
Not using Enroot? Harbor also ships a native Docker environment (
external/harbor/src/harbor/environments/docker/). To switch, replace theEnrootEnvironmentimport + instantiation insrc/memory_agent/runner.py(around lines 16 and 97) with Harbor'sDockerEnvironment, and drop theenroot:block from the YAML configs (or comment it out — theSQSH_DIR/DATA_DIRenv vars become unused).
Environment:
export OPENROUTER_API_KEY=sk-or-...
export SQSH_DIR=/path/to/enroot-cache # pre-built .sqsh task images (Enroot only)
export DATA_DIR=/path/to/enroot-data # writable enroot data root (Enroot only)
Memory-enabled run (Sonnet 4.5 action + Opus 4.6 memory):
python -m memory_agent.cli run configs/memory_terminalbench.yaml
Baseline (Sonnet 4.5, no memory):
python -m memory_agent.cli run configs/baseline_terminalbench.yaml
Run a single task (smoke test):
python -m memory_agent.cli run configs/baseline_terminalbench.yaml --task chess-best-move
Resume an interrupted run:
python -m memory_agent.cli run configs/memory_terminalbench.yaml \
--resume-run-dir ./outputs/memory-terminalbench/run_<id>
To evaluate other action / memory model pairs, copy either YAML and edit
model.model_name (action) or memory.model_name (memory). Both fields
accept any LiteLLM-compatible name; point api_base at a local vLLM
endpoint for open-weight models.
src/memory_agent/
├── memory/ # two-phase memory core (bank, trigger, prompts)
├── memory_enabled_agent.py # Terminal-Bench Terminus2 subclass with memory
├── runner.py # TB batch runner
├── cli.py # `memory-agent run …` entrypoint
├── config.py # YAML → pydantic
└── data/ # harbor + parquet task adapters
external/harbor/ # vendored Harbor fork (Enroot + Docker environments)
configs/ # TB inference YAMLs (memory + baseline)
If you use MemoryAgent, please cite our paper:
@article{wu2026memoryagent,
title = {Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents},
author = {Wu, Yifan and Zhang, Lizhu and Zhou, Yuhang and Wang, Mingyi and Peng, Bo and Li, Serena and Fan, Xiangjun and Zhao, Zhuokai},
year = {2026},
journal = {arXiv preprint arXiv:2607.08716},
url = {https://arxiv.org/abs/2607.08716}
}
Apache-2.0. See LICENSE. external/harbor/ is derived from
https://github.com/laude-institute/harbor.
1 commits
Python
100.0%
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
Python
56
1 commits
updated Jul 14, 2026
As a trajectory grows, the state that should drive the next action — requirements, environment facts, prior attempts, diagnoses, open subgoals — gets buried in the context window or pushed past it, and stops influencing decisions when it's needed. We call this behavioral state decay, and treat memory as an active intervention rather than passive retrieval. Proactive Memory Agent runs a plug-and-play memory agent alongside an unmodified action agent: it reads the recent trajectory, manages a structured memory bank, and decides whether to inject a memory-grounded reminder to the action agent or stay silent.
Rendered agent trajectories (open in browser via raw.githack.com):
| Task | Baseline | Memory (action agent trajectory) | Memory (memory agent trajectory) |
|---|---|---|---|
| git-multibranch | action agent trajectory | action agent trajectory | memory agent trajectory |
| regex-log | action agent trajectory | action agent trajectory | memory agent trajectory |
| sqlite-with-gcov | action agent trajectory | action agent trajectory | memory agent trajectory |
| hf-model-inference | action agent trajectory | action agent trajectory | memory agent trajectory |
| adaptive-rejection-sampler | action agent trajectory | action agent trajectory | memory agent trajectory |
System integration. The action agent (left) interacts with the environment; the memory agent (right) runs alongside, observing a sliding window of recent steps and the current memory store. At every N steps the memory agent is invoked to update the bank and optionally inject a context reminder into the next action-agent call.

Memory agent internals. Phase 1 manages the memory bank (status, knowledge,
procedural entries) through tool calls. Phase 2 reads the updated bank and either
emits a <context_for_action> reminder or <no_intervention/>.

git clone https://github.com/yifannnwu/proactive-memory-agent.git
cd proactive-memory-agent
# 1. Install the vendored Harbor (required by the Terminal-Bench runner)
pip install -e external/harbor
# 2. Install this package
pip install -e .
Python 3.12+. Uses litellm for LLM calls (works with OpenRouter, OpenAI,
Anthropic, local vLLM, etc.).
Quick import check:
python -c "import memory_agent; print('OK')"
python -m memory_agent.cli --help
The Terminal-Bench runner uses Harbor + Enroot for sandboxed shell execution
(chosen because our cluster is rootless and shares an FSx cache of pre-built
.sqsh images). Both configs run the full terminal-bench@2.0 set (89 tasks).
Not using Enroot? Harbor also ships a native Docker environment (
external/harbor/src/harbor/environments/docker/). To switch, replace theEnrootEnvironmentimport + instantiation insrc/memory_agent/runner.py(around lines 16 and 97) with Harbor'sDockerEnvironment, and drop theenroot:block from the YAML configs (or comment it out — theSQSH_DIR/DATA_DIRenv vars become unused).
Environment:
export OPENROUTER_API_KEY=sk-or-...
export SQSH_DIR=/path/to/enroot-cache # pre-built .sqsh task images (Enroot only)
export DATA_DIR=/path/to/enroot-data # writable enroot data root (Enroot only)
Memory-enabled run (Sonnet 4.5 action + Opus 4.6 memory):
python -m memory_agent.cli run configs/memory_terminalbench.yaml
Baseline (Sonnet 4.5, no memory):
python -m memory_agent.cli run configs/baseline_terminalbench.yaml
Run a single task (smoke test):
python -m memory_agent.cli run configs/baseline_terminalbench.yaml --task chess-best-move
Resume an interrupted run:
python -m memory_agent.cli run configs/memory_terminalbench.yaml \
--resume-run-dir ./outputs/memory-terminalbench/run_<id>
To evaluate other action / memory model pairs, copy either YAML and edit
model.model_name (action) or memory.model_name (memory). Both fields
accept any LiteLLM-compatible name; point api_base at a local vLLM
endpoint for open-weight models.
src/memory_agent/
├── memory/ # two-phase memory core (bank, trigger, prompts)
├── memory_enabled_agent.py # Terminal-Bench Terminus2 subclass with memory
├── runner.py # TB batch runner
├── cli.py # `memory-agent run …` entrypoint
├── config.py # YAML → pydantic
└── data/ # harbor + parquet task adapters
external/harbor/ # vendored Harbor fork (Enroot + Docker environments)
configs/ # TB inference YAMLs (memory + baseline)
If you use MemoryAgent, please cite our paper:
@article{wu2026memoryagent,
title = {Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents},
author = {Wu, Yifan and Zhang, Lizhu and Zhou, Yuhang and Wang, Mingyi and Peng, Bo and Li, Serena and Fan, Xiangjun and Zhao, Zhuokai},
year = {2026},
journal = {arXiv preprint arXiv:2607.08716},
url = {https://arxiv.org/abs/2607.08716}
}
Apache-2.0. See LICENSE. external/harbor/ is derived from
https://github.com/laude-institute/harbor.
1 commits
Python
100.0%