YerbaPage/Awesome-Agent-Context-Compression

Context Compression in Long-Horizon LLM-based Agents

90

46 commits

updated Aug 26, 2026

See the code

README

Awesome Agent Context Compression

Paper Stars Awesome License: MIT PRs Welcome

A comprehensive survey and curated list of resources on Context Compression in Long-Horizon LLM-based Agents โ€” covering observation compression, trajectory compression, plan & reasoning compression, memory state compression, and representation-level compression for coding agents, web/GUI agents, research agents, and multi-agent systems.


๐Ÿ“ฐ News


๐ŸŽฏ Introduction

Introduction

With the rapid evolution of LLM agents, long context has become a central challenge across open-ended domains such as automated software engineering, visual GUI navigation, and deep research. As agents continuously interact with dynamic environments, their operational history forms an unbounded agentic trajectory. This trajectory, typified by the interleaved and heterogeneous ReAct paradigm of Actions, Thoughts, and Observations (A-T-O), can quickly exhaust the LLM context window and trigger severe context explosion. Such explosion leads to cascading failures in which information density drops, critical constraints fade, and long-horizon planning deteriorates.

  • Dynamic growth: Context expands with every observation, action, and tool output
  • Heterogeneous composition: A-T-O trajectories mix code, HTML, plans, and dialogue history
  • Multi-step dependency: Information irrelevant now may matter later
  • Error propagation: Compression mistakes compound over long horizons

This repository organizes the literature along three axes: what is selected for compression, how it is transformed, and who decides when compression occurs. These axes map onto the pipeline: targets define the input to Select (S), mechanisms implement Compress (ฮฆ) and Store (M), and control policies govern when these stages, and when necessary Recover (R), are invoked.

AxisDimensionRange
WhatCompression targetsObservation โ†’ Trajectory โ†’ Plan and reasoning โ†’ Memory state โ†’ Representation-level
HowCompression mechanismsMasking and truncation โ†’ Summarization and abstraction โ†’ Pruning and reduction โ†’ Externalization and retrieval โ†’ Representation compression
Who/WhenControl policies and intervention timingSystem-controlled โ†’ External controller โ†’ Agent-controlled โ†’ Learned

Agent context compression taxonomy


๐Ÿ“š Table of Contents


Agent Surveys

  • The Rise and Potential of Large Language Model Based Agents: A Survey, Xi et al., arXiv Badge
  • A Survey on Large Language Model Based Autonomous Agents, Wang et al., FCS Badge
  • A Survey of Context Engineering for Large Language Models, Mei et al., arXiv Badge

Prompt & Context Compression

  • LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Jiang et al., EMNLP Badge GitHub stars
  • LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Jiang et al., ACL Badge GitHub stars
  • Compressing Context to Enhance Inference Efficiency of Large Language Models, Li et al., EMNLP Badge

Tool Learning & Agent Evaluation

  • Tool Learning with Foundation Models, Qin et al., ACM Badge
  • AgentBench: Evaluating LLMs as Agents, Liu et al., ICLR Badge
  • WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al., ICLR Badge

Long Context

  • Lost in the Middle: How Language Models Use Long Contexts, Liu et al., TACL Badge
  • Longformer: The Long-Document Transformer, Beltagy et al., arXiv Badge
  • LLM Maybe LongLM: SelfExtend LLM Context Window Without Tuning, Jin et al., ICML Badge GitHub stars

๐Ÿ—๏ธ Background & Foundations

Agent context differs fundamentally from static prompts. The unified pipeline formulation is S โ†’ C โ†’ M โ†’ R (Sense โ†’ Compress โ†’ Memorize โ†’ Respond). Key challenges include:

  • Dynamic growth: Context expands with every observation, action, and tool output
  • Heterogeneous composition: A-T-O trajectories mix code, HTML, plans, and dialogue history
  • Multi-step dependency: Information irrelevant now may matter later
  • Error propagation: Compression mistakes compound over long horizons
  • ReAct: Synergizing Reasoning and Acting in Language Models, Yao et al., ICLR Badge GitHub stars

๐ŸŽฏ Compression Targets(What)

A comparative anatomy of compression targets

Observation Compression

Compressing raw environment observations (HTML pages, code files, tool outputs, screenshots) before they enter the agent context.

  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (HTML-T5), Gur et al., ICLR Badge
  • Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al., arXiv Badge
  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers, Wang et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation, Fang et al., arXiv Badge
  • LongCodeZip: Compress Long Context for Code Language Models, Shi et al., ASE Badge
  • A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al., arXiv Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al., arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge

Trajectory Compression

Compressing the accumulated actionโ€“observation history of agent execution traces.

  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al., arXiv Badge
  • AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al., ICLR Badge
  • Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma, arXiv Badge
  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al., arXiv Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al., FSE Badge
  • LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al., arXiv Badge
  • Scaling Long-Horizon LLM Agent via Context-Folding (Context-Folding), Sun et al., arXiv Badge
  • Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Plan and Reasoning Compression

Compressing planning traces, chain-of-thought reasoning, and intermediate deliberation.

  • ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al., ICLR Badge
  • PAACE: A Plan-Aware Automated Agent Context Engineering Framework, Yuksel, arXiv Badge
  • Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning (SCMA), Chen et al., arXiv Badge
  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • Compressed Step Information Memory for End-to-End Agent Foundation Models (CSIM), Liu et al., OpenReview Badge
  • SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al., arXiv Badge
  • HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (HIAGENT), Hu et al., ACL Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge

Memory State Compression

Compressing and managing long-term memory states, knowledge stores, and persistent agent state.

  • MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al., NeurIPS Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • AI Agents Need Memory Control Over More Context, Bousetouane, arXiv Badge
  • Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management, Li et al., arXiv Badge
  • A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al., arXiv Badge
  • ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents, Wu et al., arXiv Badge
  • Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (ProMem), Yang et al., arXiv Badge
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al., arXiv Badge
  • AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al., arXiv Badge
  • Git Context Controller: Manage the Context of LLM-based Agents like Git (GCC), Wu et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Representation-Level Compression

Compressing context at the embedding or KV-cache level rather than at the text level.

  • AgentOCR: Reimagining Agent History via Optical Self-Compression, Feng et al., arXiv Badge
  • Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng, ICMSCI Badge
  • Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al., arXiv Badge
  • Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors (SAC), Liu et al., arXiv Badge
  • PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction, Zhao et al., arXiv Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge

๐Ÿ”ง Compression Mechanisms(How)

Masking and Truncation

Simple but effective strategies that remove or mask parts of the context based on rules or heuristics.

  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • LLM Maybe LongLM: SelfExtend LLM Context Window Without Tuning, Jin et al., ICML Badge GitHub stars
  • Longformer: The Long-Document Transformer, Beltagy et al., arXiv Badge
  • SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al., arXiv Badge

Summarization and Abstraction

Using LLMs or specialized models to produce condensed summaries of context segments.

  • ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al., ICLR Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al., arXiv Badge
  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Jiang et al., EMNLP Badge GitHub stars
  • LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Jiang et al., ACL Badge GitHub stars
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al., FSE Badge
  • Scaling Long-Horizon LLM Agent via Context-Folding (Context-Folding), Sun et al., arXiv Badge
  • HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (HIAGENT), Hu et al., ACL Badge
  • Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (ProMem), Yang et al., arXiv Badge
  • Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Pruning and Reduction

Selectively removing less important tokens, segments, or episodes from context.

  • WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al., arXiv Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction, Zhao et al., arXiv Badge
  • Compressing Context to Enhance Inference Efficiency of Large Language Models (Selective Context), Li et al., EMNLP Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation, Fang et al., arXiv Badge
  • LongCodeZip: Compress Long Context for Code Language Models, Shi et al., ASE Badge
  • ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents, Wu et al., arXiv Badge
  • AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al., arXiv Badge
  • A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al., arXiv Badge
  • LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al., arXiv Badge
  • PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge

Externalization and Retrieval

Moving information out of the prompt into external stores and retrieving on demand.

  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research, Li et al., arXiv Badge
  • MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al., NeurIPS Badge
  • A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al., arXiv Badge
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al., arXiv Badge
  • Git Context Controller: Manage the Context of LLM-based Agents like Git (GCC), Wu et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Representation Compression

Compressing at the embedding, KV-cache, or visual representation level.

  • AgentOCR: Reimagining Agent History via Optical Self-Compression, Feng et al., arXiv Badge
  • Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng, ICMSCI Badge
  • Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al., arXiv Badge
  • Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors (SAC), Liu et al., arXiv Badge
  • Compressed Step Information Memory for End-to-End Agent Foundation Models (CSIM), Liu et al., OpenReview Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge

๐Ÿงญ Control Policies and Intervention Timing(Who/When)

Who decides when compression happens, and when should the system intervene?

This axis captures the control logic that schedules compression, recovery, or memory updates. In practice, the same compression operator can be triggered by fixed system rules, an external manager, the agent itself, or a learned policy. The key distinction is not only the trigger source, but also whether intervention is reactive, periodic, or proactive.

PolicyWho decidesTypical triggerStrengthLimitation
System-controlledFixed system ruleToken budget, step count, length thresholdCheap, predictable, easy to benchmarkSemantically blind
External controllerSeparate module / plannerUtility estimate, state monitor, retriever signalModular and stableExtra overhead and latency
Agent-controlledThe agent itselfSelf-assessed need, task state, uncertaintySemantically aware and proactiveVulnerable to self-assessment errors
LearnedTrained policyReward / utility maximizationAdaptable to task objectivesData- and compute-intensive

System-Controlled Policies

System-controlled policies apply fixed rules to trigger compression:

  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al., arXiv Badge
  • PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al., arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

External Controller Policies

External-controller policies delegate the decision to a separate module:

  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge

Agent-Controlled Policies

Agent-controlled policies make compression part of the agent's own action space:

  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma, arXiv Badge
  • AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al., ICLR Badge
  • Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management, Li et al., arXiv Badge
  • MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al., NeurIPS Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge

Learned Policies

Learned policies optimize compression behavior from data:

  • ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al., ICLR Badge
  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge

โš ๏ธ Failure Modes

Failure modes taxonomy

Failure modes cases

We organize context compression failures by the earliest stage at which they arise: F1: Pre-compression Decision Error (wrong moment, target, or granularity), F2: In-compression Information Loss (semantic or structural corruption during transformation), and F3: Post-compression Access Failure (information cannot be correctly recovered when needed).

  • F1: Pre-compression Decision Error: The system compresses too early, selects the wrong content, or uses an overly coarse granularity, causing important information to disappear before compression begins. Representative work includes Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (Yang et al., arXiv) and Context as a Tool: Context Management for Long-Horizon SWE-Agents (Liu et al., arXiv).
  • F2: In-compression Information Loss: The compression step itself distorts semantics, structure, relations, or constraints, so the compressed state is no longer faithful to the original task evidence. Representative work includes HaluMem: Evaluating Hallucinations in Memory Systems of Agents (Chen et al., arXiv), Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory (Liang et al., arXiv), and ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents (Wu et al., arXiv).
  • F3: Post-compression Access Failure: The compressed information remains stored somewhere, but retrieval or reconstruction fails later, so the agent cannot recover the right state when it is needed. Representative work includes OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory (Li et al., arXiv) and KVCache-Centric Memory for LLM Agents (Zeng et al., OpenReview).

Together, these three categories form a temporal failure taxonomy over the compression pipeline: the earliest causal failure determines the label, because downstream errors in agent workflows often propagate from an upstream mistake.


๐ŸŒ Domain-Specific Analysis

Domain specific

Domain taxonomy

Coding Agents

Pipeline-level comparison of recent production-grade coding agents

Coding agents require high structural fidelity โ€” compressed context must preserve code structure, file relationships, and error traces faithfully.

  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma, arXiv Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al., arXiv Badge
  • SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al., arXiv Badge
  • A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al., arXiv Badge
  • LongCodeZip: Compress Long Context for Code Language Models, Shi et al., ASE Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge

Web & GUI Agents

Web agents face heterogeneous observations โ€” HTML, screenshots, and DOM trees that need domain-specific compression.

  • AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al., ICLR Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al., arXiv Badge
  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al., arXiv Badge
  • A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (HTML-T5), Gur et al., ICLR Badge
  • WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al., ICLR Badge
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al., arXiv Badge
  • AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge

Research & Deep-Search Agents

Research agents need high recoverability โ€” the ability to retrieve externalized information accurately over long horizons.

  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research, Li et al., arXiv Badge
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al., arXiv Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al., FSE Badge
  • LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge

Multi-Agent Systems

Multi-agent settings face unique challenges: inter-agent communication bandwidth, shared memory compression, and coordination overhead.

  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng, ICMSCI Badge
  • Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al., arXiv Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning (SCMA), Chen et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

๐Ÿ“Š Evaluation & Benchmarks

Our survey proposes a four-dimensional evaluation metric system Q = (D, R, P, O):

DimensionDescription
D (Density)Information density after compression
R (Recoverability)Ability to retrieve externalized information
P (Error Propagation)How compression errors compound over steps
O (Overhead)Computational cost of the compression itself

Relevant Benchmarks and Evaluation

  • WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al., ICLR Badge
  • AgentBench: Evaluating LLMs as Agents, Liu et al., ICLR Badge
  • MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers, Wang et al., arXiv Badge

๐Ÿ”ฎ Future Directions

  1. Compressionโ€“Retrieval Boundary โ€” When to compress in-context vs. externalize and retrieve?
  2. Recoverable Compression โ€” Compression with guaranteed information recoverability
  3. Multi-Agent Compression โ€” Efficient shared context compression across agent teams
  4. End-to-End Training โ€” Learning compression policies jointly with agent objectives
  5. Domain-Specific Methods โ€” Tailored compression for code, web, research domains
  6. Standardized Benchmarks โ€” Unified evaluation for agent context compression

๐Ÿค Contributing

We welcome contributions! Please follow these guidelines:

  1. Fork the repository
  2. Create a feature branch
  3. Add relevant papers with proper formatting
  4. Submit a pull request with a clear description

Paper Formatting Guidelines

<li><i><b>Paper Title</b></i>, Author et al., <a href="URL" target="_blank"><img src="https://img.shields.io/badge/SOURCE-YEAR.MM-COLOR" alt="SOURCE Badge"></a></li>

Badge Colors

  • arXiv Badge red for arXiv papers
  • Conference Badge blue for conference/journal papers
  • OpenReview Badge orange for OpenReview submissions
  • GitHub Badge white for GitHub repositories

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.


๐Ÿ“‘ Citation

If you find this survey helpful in your research, please consider citing:

@article{202605.2065,
	doi = {10.20944/preprints202605.2065.v1},
	url = {https://doi.org/10.20944/preprints202605.2065.v1},
	year = 2026,
	month = {May},
	publisher = {Preprints},
	author = {Yifei Wang and Ziteng Wang and Yuling Shi and Silin Chen and Xinrui Wang and Yueqi Wang and Beijun Shen and Linjing Li and Xiaodong Gu and Julian McAuley and Daniel Dajun Zeng},
	title = {Context Compression for LLM Agents: A Survey of Methods, Failure Modes, and Evaluation},
	journal = {Preprints}
}

Star โญ this repository if you find it helpful!

This repository is actively maintained and we keep updating it with the latest work on context compression for LLM agents. Contributions are very welcome โ€” feel free to open an issue or pull request!


โญ Star History

Star History Chart


Contributors

YerbaPage

6 commits

cslsolow

3 commits

YerbaPage/Awesome-Agent-Context-Compression

Context Compression in Long-Horizon LLM-based Agents

90

46 commits

updated Aug 26, 2026

See the code

README

Awesome Agent Context Compression

Paper Stars Awesome License: MIT PRs Welcome

A comprehensive survey and curated list of resources on Context Compression in Long-Horizon LLM-based Agents โ€” covering observation compression, trajectory compression, plan & reasoning compression, memory state compression, and representation-level compression for coding agents, web/GUI agents, research agents, and multi-agent systems.


๐Ÿ“ฐ News


๐ŸŽฏ Introduction

Introduction

With the rapid evolution of LLM agents, long context has become a central challenge across open-ended domains such as automated software engineering, visual GUI navigation, and deep research. As agents continuously interact with dynamic environments, their operational history forms an unbounded agentic trajectory. This trajectory, typified by the interleaved and heterogeneous ReAct paradigm of Actions, Thoughts, and Observations (A-T-O), can quickly exhaust the LLM context window and trigger severe context explosion. Such explosion leads to cascading failures in which information density drops, critical constraints fade, and long-horizon planning deteriorates.

  • Dynamic growth: Context expands with every observation, action, and tool output
  • Heterogeneous composition: A-T-O trajectories mix code, HTML, plans, and dialogue history
  • Multi-step dependency: Information irrelevant now may matter later
  • Error propagation: Compression mistakes compound over long horizons

This repository organizes the literature along three axes: what is selected for compression, how it is transformed, and who decides when compression occurs. These axes map onto the pipeline: targets define the input to Select (S), mechanisms implement Compress (ฮฆ) and Store (M), and control policies govern when these stages, and when necessary Recover (R), are invoked.

AxisDimensionRange
WhatCompression targetsObservation โ†’ Trajectory โ†’ Plan and reasoning โ†’ Memory state โ†’ Representation-level
HowCompression mechanismsMasking and truncation โ†’ Summarization and abstraction โ†’ Pruning and reduction โ†’ Externalization and retrieval โ†’ Representation compression
Who/WhenControl policies and intervention timingSystem-controlled โ†’ External controller โ†’ Agent-controlled โ†’ Learned

Agent context compression taxonomy


๐Ÿ“š Table of Contents


Agent Surveys

  • The Rise and Potential of Large Language Model Based Agents: A Survey, Xi et al., arXiv Badge
  • A Survey on Large Language Model Based Autonomous Agents, Wang et al., FCS Badge
  • A Survey of Context Engineering for Large Language Models, Mei et al., arXiv Badge

Prompt & Context Compression

  • LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Jiang et al., EMNLP Badge GitHub stars
  • LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Jiang et al., ACL Badge GitHub stars
  • Compressing Context to Enhance Inference Efficiency of Large Language Models, Li et al., EMNLP Badge

Tool Learning & Agent Evaluation

  • Tool Learning with Foundation Models, Qin et al., ACM Badge
  • AgentBench: Evaluating LLMs as Agents, Liu et al., ICLR Badge
  • WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al., ICLR Badge

Long Context

  • Lost in the Middle: How Language Models Use Long Contexts, Liu et al., TACL Badge
  • Longformer: The Long-Document Transformer, Beltagy et al., arXiv Badge
  • LLM Maybe LongLM: SelfExtend LLM Context Window Without Tuning, Jin et al., ICML Badge GitHub stars

๐Ÿ—๏ธ Background & Foundations

Agent context differs fundamentally from static prompts. The unified pipeline formulation is S โ†’ C โ†’ M โ†’ R (Sense โ†’ Compress โ†’ Memorize โ†’ Respond). Key challenges include:

  • Dynamic growth: Context expands with every observation, action, and tool output
  • Heterogeneous composition: A-T-O trajectories mix code, HTML, plans, and dialogue history
  • Multi-step dependency: Information irrelevant now may matter later
  • Error propagation: Compression mistakes compound over long horizons
  • ReAct: Synergizing Reasoning and Acting in Language Models, Yao et al., ICLR Badge GitHub stars

๐ŸŽฏ Compression Targets(What)

A comparative anatomy of compression targets

Observation Compression

Compressing raw environment observations (HTML pages, code files, tool outputs, screenshots) before they enter the agent context.

  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (HTML-T5), Gur et al., ICLR Badge
  • Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al., arXiv Badge
  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers, Wang et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation, Fang et al., arXiv Badge
  • LongCodeZip: Compress Long Context for Code Language Models, Shi et al., ASE Badge
  • A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al., arXiv Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al., arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge

Trajectory Compression

Compressing the accumulated actionโ€“observation history of agent execution traces.

  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al., arXiv Badge
  • AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al., ICLR Badge
  • Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma, arXiv Badge
  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al., arXiv Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al., FSE Badge
  • LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al., arXiv Badge
  • Scaling Long-Horizon LLM Agent via Context-Folding (Context-Folding), Sun et al., arXiv Badge
  • Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Plan and Reasoning Compression

Compressing planning traces, chain-of-thought reasoning, and intermediate deliberation.

  • ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al., ICLR Badge
  • PAACE: A Plan-Aware Automated Agent Context Engineering Framework, Yuksel, arXiv Badge
  • Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning (SCMA), Chen et al., arXiv Badge
  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • Compressed Step Information Memory for End-to-End Agent Foundation Models (CSIM), Liu et al., OpenReview Badge
  • SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al., arXiv Badge
  • HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (HIAGENT), Hu et al., ACL Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge

Memory State Compression

Compressing and managing long-term memory states, knowledge stores, and persistent agent state.

  • MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al., NeurIPS Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • AI Agents Need Memory Control Over More Context, Bousetouane, arXiv Badge
  • Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management, Li et al., arXiv Badge
  • A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al., arXiv Badge
  • ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents, Wu et al., arXiv Badge
  • Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (ProMem), Yang et al., arXiv Badge
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al., arXiv Badge
  • AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al., arXiv Badge
  • Git Context Controller: Manage the Context of LLM-based Agents like Git (GCC), Wu et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Representation-Level Compression

Compressing context at the embedding or KV-cache level rather than at the text level.

  • AgentOCR: Reimagining Agent History via Optical Self-Compression, Feng et al., arXiv Badge
  • Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng, ICMSCI Badge
  • Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al., arXiv Badge
  • Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors (SAC), Liu et al., arXiv Badge
  • PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction, Zhao et al., arXiv Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge

๐Ÿ”ง Compression Mechanisms(How)

Masking and Truncation

Simple but effective strategies that remove or mask parts of the context based on rules or heuristics.

  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • LLM Maybe LongLM: SelfExtend LLM Context Window Without Tuning, Jin et al., ICML Badge GitHub stars
  • Longformer: The Long-Document Transformer, Beltagy et al., arXiv Badge
  • SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al., arXiv Badge

Summarization and Abstraction

Using LLMs or specialized models to produce condensed summaries of context segments.

  • ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al., ICLR Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al., arXiv Badge
  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Jiang et al., EMNLP Badge GitHub stars
  • LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Jiang et al., ACL Badge GitHub stars
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al., FSE Badge
  • Scaling Long-Horizon LLM Agent via Context-Folding (Context-Folding), Sun et al., arXiv Badge
  • HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (HIAGENT), Hu et al., ACL Badge
  • Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (ProMem), Yang et al., arXiv Badge
  • Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Pruning and Reduction

Selectively removing less important tokens, segments, or episodes from context.

  • WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al., arXiv Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction, Zhao et al., arXiv Badge
  • Compressing Context to Enhance Inference Efficiency of Large Language Models (Selective Context), Li et al., EMNLP Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation, Fang et al., arXiv Badge
  • LongCodeZip: Compress Long Context for Code Language Models, Shi et al., ASE Badge
  • ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents, Wu et al., arXiv Badge
  • AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al., arXiv Badge
  • A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al., arXiv Badge
  • LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al., arXiv Badge
  • PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge

Externalization and Retrieval

Moving information out of the prompt into external stores and retrieving on demand.

  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research, Li et al., arXiv Badge
  • MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al., NeurIPS Badge
  • A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al., arXiv Badge
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al., arXiv Badge
  • Git Context Controller: Manage the Context of LLM-based Agents like Git (GCC), Wu et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

Representation Compression

Compressing at the embedding, KV-cache, or visual representation level.

  • AgentOCR: Reimagining Agent History via Optical Self-Compression, Feng et al., arXiv Badge
  • Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng, ICMSCI Badge
  • Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al., arXiv Badge
  • Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors (SAC), Liu et al., arXiv Badge
  • Compressed Step Information Memory for End-to-End Agent Foundation Models (CSIM), Liu et al., OpenReview Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge

๐Ÿงญ Control Policies and Intervention Timing(Who/When)

Who decides when compression happens, and when should the system intervene?

This axis captures the control logic that schedules compression, recovery, or memory updates. In practice, the same compression operator can be triggered by fixed system rules, an external manager, the agent itself, or a learned policy. The key distinction is not only the trigger source, but also whether intervention is reactive, periodic, or proactive.

PolicyWho decidesTypical triggerStrengthLimitation
System-controlledFixed system ruleToken budget, step count, length thresholdCheap, predictable, easy to benchmarkSemantically blind
External controllerSeparate module / plannerUtility estimate, state monitor, retriever signalModular and stableExtra overhead and latency
Agent-controlledThe agent itselfSelf-assessed need, task state, uncertaintySemantically aware and proactiveVulnerable to self-assessment errors
LearnedTrained policyReward / utility maximizationAdaptable to task objectivesData- and compute-intensive

System-Controlled Policies

System-controlled policies apply fixed rules to trigger compression:

  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability, Min et al., arXiv Badge
  • PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents, Ghulyani et al., arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents, Huang et al., arXiv Badge
  • GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression, Miao et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

External Controller Policies

External-controller policies delegate the decision to a separate module:

  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems, Dadhich, arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression, Trukhina et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge

Agent-Controlled Policies

Agent-controlled policies make compression part of the agent's own action space:

  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma, arXiv Badge
  • AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al., ICLR Badge
  • Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management, Li et al., arXiv Badge
  • MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents, Zhou et al., NeurIPS Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge

Learned Policies

Learned policies optimize compression behavior from data:

  • ACON: Optimizing Context Compression for Long-horizon LLM Agents, Kang et al., ICLR Badge
  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context-Driven Incremental Compression for Multi-Turn Dialogue Generation, Jung et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven, Tang et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory, Yang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge
  • SeDeM: Selective Decompression of Hidden-State Memories for Long-Context QA, Haghifam et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge

โš ๏ธ Failure Modes

Failure modes taxonomy

Failure modes cases

We organize context compression failures by the earliest stage at which they arise: F1: Pre-compression Decision Error (wrong moment, target, or granularity), F2: In-compression Information Loss (semantic or structural corruption during transformation), and F3: Post-compression Access Failure (information cannot be correctly recovered when needed).

  • F1: Pre-compression Decision Error: The system compresses too early, selects the wrong content, or uses an overly coarse granularity, causing important information to disappear before compression begins. Representative work includes Beyond Static Summarization: Proactive Memory Extraction for LLM Agents (Yang et al., arXiv) and Context as a Tool: Context Management for Long-Horizon SWE-Agents (Liu et al., arXiv).
  • F2: In-compression Information Loss: The compression step itself distorts semantics, structure, relations, or constraints, so the compressed state is no longer faithful to the original task evidence. Representative work includes HaluMem: Evaluating Hallucinations in Memory Systems of Agents (Chen et al., arXiv), Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory (Liang et al., arXiv), and ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents (Wu et al., arXiv).
  • F3: Post-compression Access Failure: The compressed information remains stored somewhere, but retrieval or reconstruction fails later, so the agent cannot recover the right state when it is needed. Representative work includes OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory (Li et al., arXiv) and KVCache-Centric Memory for LLM Agents (Zeng et al., OpenReview).

Together, these three categories form a temporal failure taxonomy over the compression pipeline: the earliest causal failure determines the label, because downstream errors in agent workflows often propagate from an upstream mistake.


๐ŸŒ Domain-Specific Analysis

Domain specific

Domain taxonomy

Coding Agents

Pipeline-level comparison of recent production-grade coding agents

Coding agents require high structural fidelity โ€” compressed context must preserve code structure, file relationships, and error traces faithfully.

  • SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Wang et al., arXiv Badge
  • Context as a Tool: Context Management for Long-Horizon SWE-Agents (CAT), Liu et al., arXiv Badge
  • The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management, Lindenbauer et al., NeurIPS Workshop Badge
  • Active Context Compression: Autonomous Memory Management in LLM Agents (Focus/ACC), Verma, arXiv Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (AgentDiet), Xiao et al., FSE Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management (LoCoEval Framework), Liu et al., arXiv Badge
  • SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context, Lian et al., arXiv Badge
  • A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression (TACO), Ren et al., arXiv Badge
  • LongCodeZip: Compress Long Context for Code Language Models, Shi et al., ASE Badge
  • CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding, Shi et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • CoMem: Context Management with a Decoupled Long-Context Model, Zhang et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge
  • MESA: Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory, Zhao et al., arXiv Badge

Web & GUI Agents

Web agents face heterogeneous observations โ€” HTML, screenshots, and DOM trees that need domain-specific compression.

  • AgentFold: Long-Horizon Web Agents with Proactive Context Management, Ye et al., ICLR Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents (LCoW), Lee et al., arXiv Badge
  • PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents, Liu et al., arXiv Badge
  • WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning, Wang et al., arXiv Badge
  • A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (HTML-T5), Gur et al., ICLR Badge
  • WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al., ICLR Badge
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory, Li et al., arXiv Badge
  • AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management, Tian et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management, Liu et al., arXiv Badge
  • Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents, Lin et al., arXiv Badge

Research & Deep-Search Agents

Research agents need high recoverability โ€” the ability to retrieve externalized information accurately over long horizons.

  • COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context, Wan et al., arXiv Badge
  • Step-DeepResearch Technical Report, Hu et al., arXiv Badge
  • WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research, Li et al., arXiv Badge
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents, Wu et al., arXiv Badge
  • ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization, Wu et al., arXiv Badge
  • Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (BACM), Xiao et al., FSE Badge
  • LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents, Lu et al., arXiv Badge
  • ACM: Agentic Context Management for Long Horizon Tasks, Li et al., arXiv Badge
  • Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget, Bei et al., arXiv Badge
  • Learning Agent-Compatible Context Management for Long-Horizon Tasks, Yi et al., arXiv Badge
  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay, Hu et al., arXiv Badge

Multi-Agent Systems

Multi-agent settings face unique challenges: inter-agent communication bandwidth, shared memory compression, and coordination overhead.

  • AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression, Bai et al., arXiv Badge
  • Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression, Kriuk & Ng, ICMSCI Badge
  • Cross-Modal Memory Compression for Efficient Multi-Agent Debate (DebateOCR), Wu et al., arXiv Badge
  • ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization, Su et al., arXiv Badge
  • Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning (SCMA), Chen et al., arXiv Badge
  • EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management, Yang et al., arXiv Badge
  • Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review, Chen et al., arXiv Badge

๐Ÿ“Š Evaluation & Benchmarks

Our survey proposes a four-dimensional evaluation metric system Q = (D, R, P, O):

DimensionDescription
D (Density)Information density after compression
R (Recoverability)Ability to retrieve externalized information
P (Error Propagation)How compression errors compound over steps
O (Overhead)Computational cost of the compression itself

Relevant Benchmarks and Evaluation

  • WebArena: A Realistic Web Environment for Building Autonomous Agents, Zhou et al., ICLR Badge
  • AgentBench: Evaluating LLMs as Agents, Liu et al., ICLR Badge
  • MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers, Wang et al., arXiv Badge

๐Ÿ”ฎ Future Directions

  1. Compressionโ€“Retrieval Boundary โ€” When to compress in-context vs. externalize and retrieve?
  2. Recoverable Compression โ€” Compression with guaranteed information recoverability
  3. Multi-Agent Compression โ€” Efficient shared context compression across agent teams
  4. End-to-End Training โ€” Learning compression policies jointly with agent objectives
  5. Domain-Specific Methods โ€” Tailored compression for code, web, research domains
  6. Standardized Benchmarks โ€” Unified evaluation for agent context compression

๐Ÿค Contributing

We welcome contributions! Please follow these guidelines:

  1. Fork the repository
  2. Create a feature branch
  3. Add relevant papers with proper formatting
  4. Submit a pull request with a clear description

Paper Formatting Guidelines

<li><i><b>Paper Title</b></i>, Author et al., <a href="URL" target="_blank"><img src="https://img.shields.io/badge/SOURCE-YEAR.MM-COLOR" alt="SOURCE Badge"></a></li>

Badge Colors

  • arXiv Badge red for arXiv papers
  • Conference Badge blue for conference/journal papers
  • OpenReview Badge orange for OpenReview submissions
  • GitHub Badge white for GitHub repositories

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.


๐Ÿ“‘ Citation

If you find this survey helpful in your research, please consider citing:

@article{202605.2065,
	doi = {10.20944/preprints202605.2065.v1},
	url = {https://doi.org/10.20944/preprints202605.2065.v1},
	year = 2026,
	month = {May},
	publisher = {Preprints},
	author = {Yifei Wang and Ziteng Wang and Yuling Shi and Silin Chen and Xinrui Wang and Yueqi Wang and Beijun Shen and Linjing Li and Xiaodong Gu and Julian McAuley and Daniel Dajun Zeng},
	title = {Context Compression for LLM Agents: A Survey of Methods, Failure Modes, and Evaluation},
	journal = {Preprints}
}

Star โญ this repository if you find it helpful!

This repository is actively maintained and we keep updating it with the latest work on context compression for LLM agents. Contributions are very welcome โ€” feel free to open an issue or pull request!


โญ Star History

Star History Chart


Contributors

YerbaPage

6 commits

cslsolow

3 commits