agentcacheback/cacheback

CacheBack: receiver-conditioned latent communication between agents

Python

6

6 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

LLM agents can communicate without words, and now without sharing their entire context. (r/LocalLLaMA)

TL;DR: What an agent sends should depend on what the next agent needs. CacheBack lets agents share a selected subset of their internal state. With Qwen3-8B on FanOutQA, it achieves 3.2× faster median task completion and 14.7 percentage points higher accuracy than same-size text communication. It’s…

3

Oct 6, 2026

README

CacheBack

Receiver-conditioned latent communication for agents.

Code for Receiver-Conditioned Latent Communication gives 94% CacheBack. Website · Paper · Python API · Paper replication

Agents distribute large contexts across senders and receivers. Text messages require decoding and can omit evidence. Full KV-cache messages accumulate every sender's context at the receiver, increasing memory use and context length. Receiver-conditioned communication lets the receiver state what it needs in a request. CacheBack is a training-free selector that uses attention to that request to choose which sender positions enter the handoff.

FanOutQA accuracy and latency across four model families, alongside receiver-conditioned selection of sender state.

Highest-accuracy CacheBack setting per family versus same-size text on FanOutQA, with 50 concurrent tasks on eight H100 GPUs. Operating points differ by family. Right: the receiver query guides which sender positions enter the handoff.

Architecture animation

▶ Explore the interactive architecture animation

CacheBack selects sender positions for the receiver's request and passes them with latent steps, keeping the receiver within its context limit.

Click the preview to open the interactive animation, then press Play. It shows how receiver-conditioned selection keeps the handoff small.

Watch the demo

▶ Watch the 36-second coding demo

CacheBack finishes a Django fix while text workers are still generating messages.

Download the updated 4K video.

Seven Qwen3-8B workers help a coordinator fix a Django bug. CacheBack completes this recorded case in 25.66 s, versus 113.21 s for text: 4.41× faster. Both produce the same patch and subsequently pass all 88 tests. Startup and test grading are excluded from the clocks. The silent video reconstructs the interface from separate runs and varies playback speed. This is one case, not an aggregate benchmark. Recorded results.

Also explore the interactive booking relay to inspect selected positions and text messages at each handoff. Its demo guide covers recording your own.

Install

Install and import the package as rclc. Until the PyPI release, install from a checkout:

python -m pip install .
python examples/quickstart.py   # Hugging Face, runs on CPU with Qwen3-0.6B

The default GPU runtime is dense Qwen3 on vLLM 0.11.1, the paper's Qwen stack. In a Linux CUDA environment with Python 3.10 or newer:

python -m pip install -r requirements/vllm.txt
python -m pip install .

The engine needs the capture connector; see the vLLM setup and integration example. vLLM 0.26.0 is an explicit opt-in. Both pins passed a Qwen3-8B GPU journey on an A100 (results).

Use it

Bind existing Hugging Face Qwen3 agents by their loaded model and tokenizer:

import rclc

agent_x = rclc.bind(model, tokenizer, messages=history_x, backend="hf")
agent_y = rclc.bind(model, tokenizer, messages=history_y, backend="hf")
await rclc.transfer(agent_x, agent_y, "Who owns Cedar?")

inputs = agent_y.pop()
output = model.generate(**inputs, return_dict_in_generate=True)
agent_y.update(inputs=inputs, generation=output)

messages is each agent's chat history; the request that guides selection goes in transfer. Either agent can send or receive, sharing one model or matching copies. Transfer never generates an answer. Cross-model alignment is not supported.

Supply state withUse when
messages=historyYou have chat messages; RCLC applies the chat template.
prompt=text_or_idsYou have rendered text or exact token IDs.
prompt=ids, past_key_values=cacheYou already have an HF cache.
generation=outputYou have a native HF generation result.
prompt=ids, request_id=capture_idYou already captured the vLLM prompt.
NeitherReceive first, or supply sender state later with update.
  • Outside an async function or notebook, use rclc.transfer_sync(...) with the same options. Native generation and update block; finish each operation before reusing the model (concurrency).
  • vLLM: rclc.bind(llm, backend="vllm", messages=history) with the engine setup. Local dense Qwen3 weights are required; hosted chat APIs are not supported.
  • await agent_y.append(new_text_or_ids) adds a tool result without losing received state; await agent_y.inputs() continues locally; save/load snapshot retained state (continuation and snapshots).
  • agent_y.inspect() shows queued handoffs; record_selection=True plus agent_y.selection_html() shows what survived compression (selection inspection).
  • python -m rclc doctor (--backend vllm for CUDA) or rclc.check(agent_y) checks setup without running a model or downloading weights.

Examples: quickstart, cached conversation, Colab GPU notebook.

Options

from functools import partial

await rclc.transfer(senders, receivers, requests)               # Relative r4 (default)
await rclc.transfer(senders, receivers, requests, ratio=16)     # Keep 1/16 of new positions
await rclc.transfer(senders, receivers, requests, budget=1024)  # Bounded message size
await rclc.transfer(senders, receivers, requests, selector=rclc.selectors.qsnap)
await rclc.transfer(senders, receivers, requests,
                   reasoning=partial(rclc.latent_mass, steps=50), budget=128)

Each argument is one item or a list: two senders, two receivers and two requests produce four deliveries of two messages each. See budgets, selectors, reasoning and representations.

Results in the paper

For Qwen 3 8B, the paper reports 55.3% strict accuracy versus 40.7% for same-size text, a 14.7 percentage-point gain, and 3.2x lower median task-completion latency, while removing 75% of sender positions. At 16x compression, CacheBack removes approximately 94% of sender positions while improving accuracy and latency over same-size text across the tested families and topologies. These are the paper's benchmark results, not performance claims for this package.

Bring your own selector or benchmark

Pass any callable as selector= and compare it with CacheBack using the small evaluation script. The paper's FanOutQA and LongBench v2 panels need separate integration. New benchmarks and negative results are welcome; see Contributing.

Branches

BranchWhat it provides
mainState-transfer API, Hugging Face/vLLM adapters, examples, development harness
paperFour-family runtimes, both benchmarks, twelve experiment configs, pinned serving stacks, data bundles

The paper-v1 tag fixes the replication snapshot; see the replication guide. For checks, tests and GPU validation, see the development guide.

License and citation

Apache-2.0. Data licenses and replication requirements are documented on the paper branch.

@misc{rossi2026cacheback,
  title  = {Receiver-Conditioned Latent Communication gives 94\% CacheBack},
  author = {Rossi, Maximillian and Raghunath, Prajwal and Xuan, Haoqing and Zhang, Yusen and Wu, Eugene},
  year   = {2026},
  eprint = {2609.32046},
  archivePrefix = {arXiv},
  url    = {https://arxiv.org/abs/2609.32046}
}

agentcacheback/cacheback

CacheBack: receiver-conditioned latent communication between agents

Python

6

6 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

LLM agents can communicate without words, and now without sharing their entire context. (r/LocalLLaMA)

TL;DR: What an agent sends should depend on what the next agent needs. CacheBack lets agents share a selected subset of their internal state. With Qwen3-8B on FanOutQA, it achieves 3.2× faster median task completion and 14.7 percentage points higher accuracy than same-size text communication. It’s…

3

Oct 6, 2026

README

CacheBack

Receiver-conditioned latent communication for agents.

Code for Receiver-Conditioned Latent Communication gives 94% CacheBack. Website · Paper · Python API · Paper replication

Agents distribute large contexts across senders and receivers. Text messages require decoding and can omit evidence. Full KV-cache messages accumulate every sender's context at the receiver, increasing memory use and context length. Receiver-conditioned communication lets the receiver state what it needs in a request. CacheBack is a training-free selector that uses attention to that request to choose which sender positions enter the handoff.

FanOutQA accuracy and latency across four model families, alongside receiver-conditioned selection of sender state.

Highest-accuracy CacheBack setting per family versus same-size text on FanOutQA, with 50 concurrent tasks on eight H100 GPUs. Operating points differ by family. Right: the receiver query guides which sender positions enter the handoff.

Architecture animation

▶ Explore the interactive architecture animation

CacheBack selects sender positions for the receiver's request and passes them with latent steps, keeping the receiver within its context limit.

Click the preview to open the interactive animation, then press Play. It shows how receiver-conditioned selection keeps the handoff small.

Watch the demo

▶ Watch the 36-second coding demo

CacheBack finishes a Django fix while text workers are still generating messages.

Download the updated 4K video.

Seven Qwen3-8B workers help a coordinator fix a Django bug. CacheBack completes this recorded case in 25.66 s, versus 113.21 s for text: 4.41× faster. Both produce the same patch and subsequently pass all 88 tests. Startup and test grading are excluded from the clocks. The silent video reconstructs the interface from separate runs and varies playback speed. This is one case, not an aggregate benchmark. Recorded results.

Also explore the interactive booking relay to inspect selected positions and text messages at each handoff. Its demo guide covers recording your own.

Install

Install and import the package as rclc. Until the PyPI release, install from a checkout:

python -m pip install .
python examples/quickstart.py   # Hugging Face, runs on CPU with Qwen3-0.6B

The default GPU runtime is dense Qwen3 on vLLM 0.11.1, the paper's Qwen stack. In a Linux CUDA environment with Python 3.10 or newer:

python -m pip install -r requirements/vllm.txt
python -m pip install .

The engine needs the capture connector; see the vLLM setup and integration example. vLLM 0.26.0 is an explicit opt-in. Both pins passed a Qwen3-8B GPU journey on an A100 (results).

Use it

Bind existing Hugging Face Qwen3 agents by their loaded model and tokenizer:

import rclc

agent_x = rclc.bind(model, tokenizer, messages=history_x, backend="hf")
agent_y = rclc.bind(model, tokenizer, messages=history_y, backend="hf")
await rclc.transfer(agent_x, agent_y, "Who owns Cedar?")

inputs = agent_y.pop()
output = model.generate(**inputs, return_dict_in_generate=True)
agent_y.update(inputs=inputs, generation=output)

messages is each agent's chat history; the request that guides selection goes in transfer. Either agent can send or receive, sharing one model or matching copies. Transfer never generates an answer. Cross-model alignment is not supported.

Supply state withUse when
messages=historyYou have chat messages; RCLC applies the chat template.
prompt=text_or_idsYou have rendered text or exact token IDs.
prompt=ids, past_key_values=cacheYou already have an HF cache.
generation=outputYou have a native HF generation result.
prompt=ids, request_id=capture_idYou already captured the vLLM prompt.
NeitherReceive first, or supply sender state later with update.
  • Outside an async function or notebook, use rclc.transfer_sync(...) with the same options. Native generation and update block; finish each operation before reusing the model (concurrency).
  • vLLM: rclc.bind(llm, backend="vllm", messages=history) with the engine setup. Local dense Qwen3 weights are required; hosted chat APIs are not supported.
  • await agent_y.append(new_text_or_ids) adds a tool result without losing received state; await agent_y.inputs() continues locally; save/load snapshot retained state (continuation and snapshots).
  • agent_y.inspect() shows queued handoffs; record_selection=True plus agent_y.selection_html() shows what survived compression (selection inspection).
  • python -m rclc doctor (--backend vllm for CUDA) or rclc.check(agent_y) checks setup without running a model or downloading weights.

Examples: quickstart, cached conversation, Colab GPU notebook.

Options

from functools import partial

await rclc.transfer(senders, receivers, requests)               # Relative r4 (default)
await rclc.transfer(senders, receivers, requests, ratio=16)     # Keep 1/16 of new positions
await rclc.transfer(senders, receivers, requests, budget=1024)  # Bounded message size
await rclc.transfer(senders, receivers, requests, selector=rclc.selectors.qsnap)
await rclc.transfer(senders, receivers, requests,
                   reasoning=partial(rclc.latent_mass, steps=50), budget=128)

Each argument is one item or a list: two senders, two receivers and two requests produce four deliveries of two messages each. See budgets, selectors, reasoning and representations.

Results in the paper

For Qwen 3 8B, the paper reports 55.3% strict accuracy versus 40.7% for same-size text, a 14.7 percentage-point gain, and 3.2x lower median task-completion latency, while removing 75% of sender positions. At 16x compression, CacheBack removes approximately 94% of sender positions while improving accuracy and latency over same-size text across the tested families and topologies. These are the paper's benchmark results, not performance claims for this package.

Bring your own selector or benchmark

Pass any callable as selector= and compare it with CacheBack using the small evaluation script. The paper's FanOutQA and LongBench v2 panels need separate integration. New benchmarks and negative results are welcome; see Contributing.

Branches

BranchWhat it provides
mainState-transfer API, Hugging Face/vLLM adapters, examples, development harness
paperFour-family runtimes, both benchmarks, twelve experiment configs, pinned serving stacks, data bundles

The paper-v1 tag fixes the replication snapshot; see the replication guide. For checks, tests and GPU validation, see the development guide.

License and citation

Apache-2.0. Data licenses and replication requirements are documented on the paper branch.

@misc{rossi2026cacheback,
  title  = {Receiver-Conditioned Latent Communication gives 94\% CacheBack},
  author = {Rossi, Maximillian and Raghunath, Prajwal and Xuan, Haoqing and Zhang, Yusen and Wu, Eugene},
  year   = {2026},
  eprint = {2609.32046},
  archivePrefix = {arXiv},
  url    = {https://arxiv.org/abs/2609.32046}
}