Awesome Agent Context Compression
A comprehensive survey and curated list of resources on Context Compression in Long-Horizon LLM-based Agents โ covering observation compression, trajectory compression, plan & reasoning compression, memory state compression, and representation-level compression for coding agents, web/GUI agents, research agents, and multi-agent systems.
๐ฐ News
๐ฏ Introduction
With the rapid evolution of LLM agents, long context has become a central challenge across open-ended domains such as automated software engineering, visual GUI navigation, and deep research. As agents continuously interact with dynamic environments, their operational history forms an unbounded agentic trajectory. This trajectory, typified by the interleaved and heterogeneous ReAct paradigm of Actions, Thoughts, and Observations (A-T-O), can quickly exhaust the LLM context window and trigger severe context explosion. Such explosion leads to cascading failures in which information density drops, critical constraints fade, and long-horizon planning deteriorates.
- Dynamic growth: Context expands with every observation, action, and tool output
- Heterogeneous composition: A-T-O trajectories mix code, HTML, plans, and dialogue history
- Multi-step dependency: Information irrelevant now may matter later
- Error propagation: Compression mistakes compound over long horizons
This repository organizes the literature along three axes: what is selected for compression, how it is transformed, and who decides when compression occurs. These axes map onto the pipeline: targets define the input to Select (S), mechanisms implement Compress (ฮฆ) and Store (M), and control policies govern when these stages, and when necessary Recover (R), are invoked.
| Axis | Dimension | Range |
|---|
| What | Compression targets | Observation โ Trajectory โ Plan and reasoning โ Memory state โ Representation-level |
| How | Compression mechanisms | Masking and truncation โ Summarization and abstraction โ Pruning and reduction โ Externalization and retrieval โ Representation compression |
| Who/When | Control policies and intervention timing | System-controlled โ External controller โ Agent-controlled โ Learned |
๐ Table of Contents
Agent Surveys
- The Rise and Potential of Large Language Model Based Agents: A Survey, Xi et al.,

- A Survey on Large Language Model Based Autonomous Agents, Wang et al.,

- A Survey of Context Engineering for Large Language Models, Mei et al.,

Prompt & Context Compression
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Jiang et al.,

- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Jiang et al.,

- Compressing Context to Enhance Inference Efficiency of Large Language Models, Li et al.,

Tool Learning & Agent Evaluation
- Tool Learning with Foundation Models, Qin et al.,

- AgentBench: Evaluating LLMs as Agents, Liu et al.,

- WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al.,

Long Context
- Lost in the Middle: How Language Models Use Long Contexts, Liu et al.,

- Longformer: The Long-Document Transformer, Beltagy et al.,

- LLM Maybe LongLM: SelfExtend LLM Context Window Without Tuning, Jin et al.,

๐๏ธ Background & Foundations
Agent context differs fundamentally from static prompts. The unified pipeline formulation is S โ C โ M โ R (Sense โ Compress โ Memorize โ Respond). Key challenges include:
- Dynamic growth: Context expands with every observation, action, and tool output
- Heterogeneous composition: A-T-O trajectories mix code, HTML, plans, and dialogue history
- Multi-step dependency: Information irrelevant now may matter later
- Error propagation: Compression mistakes compound over long horizons
- ReAct: Synergizing Reasoning and Acting in Language Models, Yao et al.,

๐ฏ Compression Targets(What)
Observation Compression
Compressing raw environment observations (HTML pages, code files, tool outputs, screenshots) before they enter the agent context.
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al.,

- A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (HTML-T5), Gur et al.,
-blue)
- Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al.,

- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al.,

- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al.,

- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers, Wang et al.,

- Step-DeepResearch Technical Report, Hu et al.,

- AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation, Fang et al.,

- LongCodeZip: Compress Long Context for Code Language Models, Shi et al.,

- A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al.,

- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al.,

- PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al.,

- Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

Trajectory Compression
Compressing the accumulated actionโobservation history of agent execution traces.
- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al.,

- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al.,

- WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al.,

- AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al.,

- Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma,

- Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al.,

- ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al.,

- ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al.,

- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al.,

- LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al.,

- Scaling Long-Horizon LLM Agent via Context-Folding (Context-Folding), Sun et al.,

- Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al.,

- Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al.,

- Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al.,

Plan and Reasoning Compression
Compressing planning traces, chain-of-thought reasoning, and intermediate deliberation.
- ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al.,

- PAACE: A Plan-Aware Automated Agent Context Engineering Framework, Yuksel,

- Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning (SCMA), Chen et al.,

- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al.,

- Compressed Step Information Memory for End-to-End Agent Foundation Models (CSIM), Liu et al.,

- SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al.,

- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (HIAGENT), Hu et al.,

- Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al.,

Memory State Compression
Compressing and managing long-term memory states, knowledge stores, and persistent agent state.
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al.,

- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al.,

- AI Agents Need Memory Control Over More Context, Bousetouane,

- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management, Li et al.,

- A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al.,

- ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents, Wu et al.,

- Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (ProMem), Yang et al.,

- OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al.,

- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al.,

- Git Context Controller: Manage the Context of LLM-based Agents like Git (GCC), Wu et al.,

- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich,

- Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al.,

- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al.,

- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al.,

- MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al.,

- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al.,

- Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al.,

Representation-Level Compression
Compressing context at the embedding or KV-cache level rather than at the text level.
- AgentOCR: Reimagining Agent History via Optical Self-Compression, Feng et al.,

- Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng,

- Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al.,

- Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors (SAC), Liu et al.,

- PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction, Zhao et al.,

- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al.,

- Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al.,

- CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al.,

- SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al.,

- GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al.,

๐ง Compression Mechanisms(How)
Masking and Truncation
Simple but effective strategies that remove or mask parts of the context based on rules or heuristics.
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al.,

- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al.,

- LLM Maybe LongLM: SelfExtend LLM Context Window Without Tuning, Jin et al.,

- Longformer: The Long-Document Transformer, Beltagy et al.,

- SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al.,

Summarization and Abstraction
Using LLMs or specialized models to produce condensed summaries of context segments.
- ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al.,

- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al.,

- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al.,

- Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al.,

- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al.,

- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al.,

- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Jiang et al.,

- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Jiang et al.,

- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al.,

- Scaling Long-Horizon LLM Agent via Context-Folding (Context-Folding), Sun et al.,

- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (HIAGENT), Hu et al.,

- Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (ProMem), Yang et al.,

- Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich,

- Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al.,

- Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

- Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al.,

- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al.,

- Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al.,

Pruning and Reduction
Selectively removing less important tokens, segments, or episodes from context.
- WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al.,

- ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al.,

- PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction, Zhao et al.,

- Compressing Context to Enhance Inference Efficiency of Large Language Models (Selective Context), Li et al.,

- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al.,

- AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation, Fang et al.,

- LongCodeZip: Compress Long Context for Code Language Models, Shi et al.,

- ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents, Wu et al.,

- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al.,

- A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al.,

- LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al.,

- PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al.,

- Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al.,

- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al.,

- MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al.,

- CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al.,

- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al.,

Externalization and Retrieval
Moving information out of the prompt into external stores and retrieving on demand.
- Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al.,

- Step-DeepResearch Technical Report, Hu et al.,

- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research, Li et al.,

- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al.,

- A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al.,

- OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al.,

- Git Context Controller: Manage the Context of LLM-based Agents like Git (GCC), Wu et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich,

- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al.,

- MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al.,

- SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al.,

- GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al.,

- Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al.,

Representation Compression
Compressing at the embedding, KV-cache, or visual representation level.
- AgentOCR: Reimagining Agent History via Optical Self-Compression, Feng et al.,

- Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng,

- Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al.,

- Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors (SAC), Liu et al.,

- Compressed Step Information Memory for End-to-End Agent Foundation Models (CSIM), Liu et al.,

- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al.,

- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al.,

- Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al.,

- SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al.,

- GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al.,

๐งญ Control Policies and Intervention Timing(Who/When)
Who decides when compression happens, and when should the system intervene?
This axis captures the control logic that schedules compression, recovery, or memory updates. In practice, the same compression operator can be triggered by fixed system rules, an external manager, the agent itself, or a learned policy. The key distinction is not only the trigger source, but also whether intervention is reactive, periodic, or proactive.
| Policy | Who decides | Typical trigger | Strength | Limitation |
|---|
| System-controlled | Fixed system rule | Token budget, step count, length threshold | Cheap, predictable, easy to benchmark | Semantically blind |
| External controller | Separate module / planner | Utility estimate, state monitor, retriever signal | Modular and stable | Extra overhead and latency |
| Agent-controlled | The agent itself | Self-assessed need, task state, uncertainty | Semantically aware and proactive | Vulnerable to self-assessment errors |
| Learned | Trained policy | Reward / utility maximization | Adaptable to task objectives | Data- and compute-intensive |
System-Controlled Policies
System-controlled policies apply fixed rules to trigger compression:
- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al.,

- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al.,

- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al.,

- Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al.,

- PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al.,

- Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al.,

- Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al.,

- CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al.,

- GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al.,

- Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al.,

External Controller Policies
External-controller policies delegate the decision to a separate module:
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al.,

- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al.,

- ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al.,

- Step-DeepResearch Technical Report, Hu et al.,

- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich,

- Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al.,

- Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al.,

- MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al.,

- SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al.,

Agent-Controlled Policies
Agent-controlled policies make compression part of the agent's own action space:
- Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al.,

- Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma,

- AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al.,

- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management, Li et al.,

- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al.,

Learned Policies
Learned policies optimize compression behavior from data:
- ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al.,

- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al.,

- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al.,

- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al.,

- Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

- Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al.,

- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al.,

- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al.,

- MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al.,

- SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al.,

- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al.,

โ ๏ธ Failure Modes
We organize context compression failures by the earliest stage at which they arise: F1: Pre-compression Decision Error (wrong moment, target, or granularity), F2: In-compression Information Loss (semantic or structural corruption during transformation), and F3: Post-compression Access Failure (information cannot be correctly recovered when needed).
- F1: Pre-compression Decision Error: The system compresses too early, selects the wrong content, or uses an overly coarse granularity, causing important information to disappear before compression begins. Representative work includes Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (Yang et al., arXiv) and Context as a Tool: Context Management for Long-Horizon SWE-Agents (Liu et al., arXiv).
- F2: In-compression Information Loss: The compression step itself distorts semantics, structure, relations, or constraints, so the compressed state is no longer faithful to the original task evidence. Representative work includes HaluMem: Evaluating Hallucinations in Memory Systems of Agents (Chen et al., arXiv), Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory (Liang et al., arXiv), and ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents (Wu et al., arXiv).
- F3: Post-compression Access Failure: The compressed information remains stored somewhere, but retrieval or reconstruction fails later, so the agent cannot recover the right state when it is needed. Representative work includes OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory (Li et al., arXiv) and KVCache-Centric Memory for LLM Agents (Zeng et al., OpenReview).
Together, these three categories form a temporal failure taxonomy over the compression pipeline: the earliest causal failure determines the label, because downstream errors in agent workflows often propagate from an upstream mistake.
๐ Domain-Specific Analysis
Coding Agents
Coding agents require high structural fidelity โ compressed context must preserve code structure, file relationships, and error traces faithfully.
- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al.,

- Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al.,

- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al.,

- Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma,

- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al.,

- ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al.,

- A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al.,

- SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al.,

- A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al.,

- LongCodeZip: Compress Long Context for Code Language Models, Shi et al.,

- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al.,

- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al.,

- MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al.,

Web & GUI Agents
Web agents face heterogeneous observations โ HTML, screenshots, and DOM trees that need domain-specific compression.
- AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al.,

- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al.,

- Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al.,

- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al.,

- WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al.,

- A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (HTML-T5), Gur et al.,
-blue)
- WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al.,

- OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al.,

- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al.,

- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al.,

Research & Deep-Search Agents
Research agents need high recoverability โ the ability to retrieve externalized information accurately over long horizons.
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al.,

- Step-DeepResearch Technical Report, Hu et al.,

- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research, Li et al.,

- ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al.,

- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al.,

- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al.,

- LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al.,

- ACM: Agentic Context Management for Long Horizon Tasks, Li et al.,

- Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al.,

- Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al.,

- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al.,

Multi-Agent Systems
Multi-agent settings face unique challenges: inter-agent communication bandwidth, shared memory compression, and coordination overhead.
- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al.,

- Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng,

- Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al.,

- ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al.,

- Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning (SCMA), Chen et al.,

- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al.,

- Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al.,

๐ Evaluation & Benchmarks
Our survey proposes a four-dimensional evaluation metric system Q = (D, R, P, O):
| Dimension | Description |
|---|
| D (Density) | Information density after compression |
| R (Recoverability) | Ability to retrieve externalized information |
| P (Error Propagation) | How compression errors compound over steps |
| O (Overhead) | Computational cost of the compression itself |
Relevant Benchmarks and Evaluation
- WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al.,

- AgentBench: Evaluating LLMs as Agents, Liu et al.,

- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers, Wang et al.,

๐ฎ Future Directions
- CompressionโRetrieval Boundary โ When to compress in-context vs. externalize and retrieve?
- Recoverable Compression โ Compression with guaranteed information recoverability
- Multi-Agent Compression โ Efficient shared context compression across agent teams
- End-to-End Training โ Learning compression policies jointly with agent objectives
- Domain-Specific Methods โ Tailored compression for code, web, research domains
- Standardized Benchmarks โ Unified evaluation for agent context compression
๐ค Contributing
We welcome contributions! Please follow these guidelines:
- Fork the repository
- Create a feature branch
- Add relevant papers with proper formatting
- Submit a pull request with a clear description
<li><i><b>Paper Title</b></i>, Author et al., <a href="URL" target="_blank"><img src="https://img.shields.io/badge/SOURCE-YEAR.MM-COLOR" alt="SOURCE Badge"></a></li>
Badge Colors
red for arXiv papers
blue for conference/journal papers
orange for OpenReview submissions
white for GitHub repositories
๐ License
This project is licensed under the MIT License - see the LICENSE file for details.
๐ Citation
If you find this survey helpful in your research, please consider citing:
@article{202605.2065,
doi = {10.20944/preprints202605.2065.v1},
url = {https://doi.org/10.20944/preprints202605.2065.v1},
year = 2026,
month = {May},
publisher = {Preprints},
author = {Yifei Wang and Ziteng Wang and Yuling Shi and Silin Chen and Xinrui Wang and Yueqi Wang and Beijun Shen and Linjing Li and Xiaodong Gu and Julian McAuley and Daniel Dajun Zeng},
title = {Context Compression for LLM Agents: A Survey of Methods, Failure Modes, and Evaluation},
journal = {Preprints}
}
Star โญ this repository if you find it helpful!
This repository is actively maintained and we keep updating it with the latest work on context compression for LLM agents. Contributions are very welcome โ feel free to open an issue or pull request!
โญ Star History