CacheBack: receiver-conditioned latent communication between agents
See the codeReceiver-conditioned latent communication for agents.
Code for Receiver-Conditioned Latent Communication gives 94% CacheBack. Website · Paper · Python API · Paper replication
Agents distribute large contexts across senders and receivers. Text messages require decoding and can omit evidence. Full KV-cache messages accumulate every sender's context at the receiver, increasing memory use and context length. Receiver-conditioned communication lets the receiver state what it needs in a request. CacheBack is a training-free selector that uses attention to that request to choose which sender positions enter the handoff.

Highest-accuracy CacheBack setting per family versus same-size text on FanOutQA, with 50 concurrent tasks on eight H100 GPUs. Operating points differ by family. Right: the receiver query guides which sender positions enter the handoff.
▶ Explore the interactive architecture animation
Click the preview to open the interactive animation, then press Play. It shows how receiver-conditioned selection keeps the handoff small.
▶ Watch the 36-second coding demo
Download the updated 4K video.
Seven Qwen3-8B workers help a coordinator fix a Django bug. CacheBack completes this recorded case in 25.66 s, versus 113.21 s for text: 4.41× faster. Both produce the same patch and subsequently pass all 88 tests. Startup and test grading are excluded from the clocks. The silent video reconstructs the interface from separate runs and varies playback speed. This is one case, not an aggregate benchmark. Recorded results.
Also explore the interactive booking relay to inspect selected positions and text messages at each handoff. Its demo guide covers recording your own.
Install and import the package as rclc. Until the PyPI
release, install from a checkout:
python -m pip install .
python examples/quickstart.py # Hugging Face, runs on CPU with Qwen3-0.6B
The default GPU runtime is dense Qwen3 on vLLM 0.11.1, the paper's Qwen stack. In a Linux CUDA environment with Python 3.10 or newer:
python -m pip install -r requirements/vllm.txt
python -m pip install .
The engine needs the capture connector; see the vLLM setup and integration example. vLLM 0.26.0 is an explicit opt-in. Both pins passed a Qwen3-8B GPU journey on an A100 (results).
Bind existing Hugging Face Qwen3 agents by their loaded model and tokenizer:
import rclc
agent_x = rclc.bind(model, tokenizer, messages=history_x, backend="hf")
agent_y = rclc.bind(model, tokenizer, messages=history_y, backend="hf")
await rclc.transfer(agent_x, agent_y, "Who owns Cedar?")
inputs = agent_y.pop()
output = model.generate(**inputs, return_dict_in_generate=True)
agent_y.update(inputs=inputs, generation=output)
messages is each agent's chat history; the request that guides selection goes
in transfer. Either agent can send or receive, sharing one model or matching
copies. Transfer never generates an answer. Cross-model alignment is not
supported.
| Supply state with | Use when |
|---|---|
messages=history | You have chat messages; RCLC applies the chat template. |
prompt=text_or_ids | You have rendered text or exact token IDs. |
prompt=ids, past_key_values=cache | You already have an HF cache. |
generation=output | You have a native HF generation result. |
prompt=ids, request_id=capture_id | You already captured the vLLM prompt. |
| Neither | Receive first, or supply sender state later with update. |
rclc.transfer_sync(...) with the
same options. Native generation and update block; finish each operation
before reusing the model
(concurrency).rclc.bind(llm, backend="vllm", messages=history) with the
engine setup.
Local dense Qwen3 weights are required; hosted chat APIs are not supported.await agent_y.append(new_text_or_ids) adds a tool result without losing
received state; await agent_y.inputs() continues locally; save/load
snapshot retained state
(continuation and snapshots).agent_y.inspect() shows queued handoffs; record_selection=True plus
agent_y.selection_html() shows what survived compression
(selection inspection).python -m rclc doctor (--backend vllm for CUDA) or rclc.check(agent_y)
checks setup without running a model or downloading weights.Examples: quickstart, cached conversation, Colab GPU notebook.
from functools import partial
await rclc.transfer(senders, receivers, requests) # Relative r4 (default)
await rclc.transfer(senders, receivers, requests, ratio=16) # Keep 1/16 of new positions
await rclc.transfer(senders, receivers, requests, budget=1024) # Bounded message size
await rclc.transfer(senders, receivers, requests, selector=rclc.selectors.qsnap)
await rclc.transfer(senders, receivers, requests,
reasoning=partial(rclc.latent_mass, steps=50), budget=128)
Each argument is one item or a list: two senders, two receivers and two requests produce four deliveries of two messages each. See budgets, selectors, reasoning and representations.
For Qwen 3 8B, the paper reports 55.3% strict accuracy versus 40.7% for same-size text, a 14.7 percentage-point gain, and 3.2x lower median task-completion latency, while removing 75% of sender positions. At 16x compression, CacheBack removes approximately 94% of sender positions while improving accuracy and latency over same-size text across the tested families and topologies. These are the paper's benchmark results, not performance claims for this package.
Pass any callable as selector= and compare it with CacheBack using the
small evaluation script.
The paper's FanOutQA and LongBench v2 panels need separate integration.
New benchmarks and negative results are welcome; see
Contributing.
| Branch | What it provides |
|---|---|
main | State-transfer API, Hugging Face/vLLM adapters, examples, development harness |
paper | Four-family runtimes, both benchmarks, twelve experiment configs, pinned serving stacks, data bundles |
The paper-v1 tag fixes the replication snapshot; see the
replication guide.
For checks, tests and GPU validation, see the
development guide.
Apache-2.0. Data licenses and replication requirements are documented on the
paper branch.
@misc{rossi2026cacheback,
title = {Receiver-Conditioned Latent Communication gives 94\% CacheBack},
author = {Rossi, Maximillian and Raghunath, Prajwal and Xuan, Haoqing and Zhang, Yusen and Wu, Eugene},
year = {2026},
eprint = {2609.32046},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.32046}
}
CacheBack: receiver-conditioned latent communication between agents
See the codeReceiver-conditioned latent communication for agents.
Code for Receiver-Conditioned Latent Communication gives 94% CacheBack. Website · Paper · Python API · Paper replication
Agents distribute large contexts across senders and receivers. Text messages require decoding and can omit evidence. Full KV-cache messages accumulate every sender's context at the receiver, increasing memory use and context length. Receiver-conditioned communication lets the receiver state what it needs in a request. CacheBack is a training-free selector that uses attention to that request to choose which sender positions enter the handoff.

Highest-accuracy CacheBack setting per family versus same-size text on FanOutQA, with 50 concurrent tasks on eight H100 GPUs. Operating points differ by family. Right: the receiver query guides which sender positions enter the handoff.
▶ Explore the interactive architecture animation
Click the preview to open the interactive animation, then press Play. It shows how receiver-conditioned selection keeps the handoff small.
▶ Watch the 36-second coding demo
Download the updated 4K video.
Seven Qwen3-8B workers help a coordinator fix a Django bug. CacheBack completes this recorded case in 25.66 s, versus 113.21 s for text: 4.41× faster. Both produce the same patch and subsequently pass all 88 tests. Startup and test grading are excluded from the clocks. The silent video reconstructs the interface from separate runs and varies playback speed. This is one case, not an aggregate benchmark. Recorded results.
Also explore the interactive booking relay to inspect selected positions and text messages at each handoff. Its demo guide covers recording your own.
Install and import the package as rclc. Until the PyPI
release, install from a checkout:
python -m pip install .
python examples/quickstart.py # Hugging Face, runs on CPU with Qwen3-0.6B
The default GPU runtime is dense Qwen3 on vLLM 0.11.1, the paper's Qwen stack. In a Linux CUDA environment with Python 3.10 or newer:
python -m pip install -r requirements/vllm.txt
python -m pip install .
The engine needs the capture connector; see the vLLM setup and integration example. vLLM 0.26.0 is an explicit opt-in. Both pins passed a Qwen3-8B GPU journey on an A100 (results).
Bind existing Hugging Face Qwen3 agents by their loaded model and tokenizer:
import rclc
agent_x = rclc.bind(model, tokenizer, messages=history_x, backend="hf")
agent_y = rclc.bind(model, tokenizer, messages=history_y, backend="hf")
await rclc.transfer(agent_x, agent_y, "Who owns Cedar?")
inputs = agent_y.pop()
output = model.generate(**inputs, return_dict_in_generate=True)
agent_y.update(inputs=inputs, generation=output)
messages is each agent's chat history; the request that guides selection goes
in transfer. Either agent can send or receive, sharing one model or matching
copies. Transfer never generates an answer. Cross-model alignment is not
supported.
| Supply state with | Use when |
|---|---|
messages=history | You have chat messages; RCLC applies the chat template. |
prompt=text_or_ids | You have rendered text or exact token IDs. |
prompt=ids, past_key_values=cache | You already have an HF cache. |
generation=output | You have a native HF generation result. |
prompt=ids, request_id=capture_id | You already captured the vLLM prompt. |
| Neither | Receive first, or supply sender state later with update. |
rclc.transfer_sync(...) with the
same options. Native generation and update block; finish each operation
before reusing the model
(concurrency).rclc.bind(llm, backend="vllm", messages=history) with the
engine setup.
Local dense Qwen3 weights are required; hosted chat APIs are not supported.await agent_y.append(new_text_or_ids) adds a tool result without losing
received state; await agent_y.inputs() continues locally; save/load
snapshot retained state
(continuation and snapshots).agent_y.inspect() shows queued handoffs; record_selection=True plus
agent_y.selection_html() shows what survived compression
(selection inspection).python -m rclc doctor (--backend vllm for CUDA) or rclc.check(agent_y)
checks setup without running a model or downloading weights.Examples: quickstart, cached conversation, Colab GPU notebook.
from functools import partial
await rclc.transfer(senders, receivers, requests) # Relative r4 (default)
await rclc.transfer(senders, receivers, requests, ratio=16) # Keep 1/16 of new positions
await rclc.transfer(senders, receivers, requests, budget=1024) # Bounded message size
await rclc.transfer(senders, receivers, requests, selector=rclc.selectors.qsnap)
await rclc.transfer(senders, receivers, requests,
reasoning=partial(rclc.latent_mass, steps=50), budget=128)
Each argument is one item or a list: two senders, two receivers and two requests produce four deliveries of two messages each. See budgets, selectors, reasoning and representations.
For Qwen 3 8B, the paper reports 55.3% strict accuracy versus 40.7% for same-size text, a 14.7 percentage-point gain, and 3.2x lower median task-completion latency, while removing 75% of sender positions. At 16x compression, CacheBack removes approximately 94% of sender positions while improving accuracy and latency over same-size text across the tested families and topologies. These are the paper's benchmark results, not performance claims for this package.
Pass any callable as selector= and compare it with CacheBack using the
small evaluation script.
The paper's FanOutQA and LongBench v2 panels need separate integration.
New benchmarks and negative results are welcome; see
Contributing.
| Branch | What it provides |
|---|---|
main | State-transfer API, Hugging Face/vLLM adapters, examples, development harness |
paper | Four-family runtimes, both benchmarks, twelve experiment configs, pinned serving stacks, data bundles |
The paper-v1 tag fixes the replication snapshot; see the
replication guide.
For checks, tests and GPU validation, see the
development guide.
Apache-2.0. Data licenses and replication requirements are documented on the
paper branch.
@misc{rossi2026cacheback,
title = {Receiver-Conditioned Latent Communication gives 94\% CacheBack},
author = {Rossi, Maximillian and Raghunath, Prajwal and Xuan, Haoqing and Zhang, Yusen and Wu, Eugene},
year = {2026},
eprint = {2609.32046},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.32046}
}