Training-free LLM memory compression. 100% critical retention across our benchmarks; structural guarantee under budget pressure.
Python
7
27 commits
updated Oct 5, 2026
Compress multi-turn LLM conversations by 80%+ while every constraint and decision survives.
100% critical retention across our benchmark suite — under impossibly small budgets, criticals are trimmed to their word floor and dropped only as a documented last resort.
Long conversations eat your context window. Naive truncation drops the constraint from turn 3 that the entire system depends on. DSPM fixes this.
DSPM converts each conversation turn into typed semantic patches — constraint, decision, code, entity, structure — and compresses them under a fixed token budget. Critical patches (constraints and decisions) are structurally protected: they survive compression even when everything else is trimmed.
Raw conversation (452 tokens, 18 turns)
↓
Semantic extraction → 28 patches, 18 critical
↓
7-stage compression pipeline
↓
Compressed context (249 tokens) — 100% of criticals intact
The 7-stage pipeline: dedup → slot fusion → delta encoding → causal pruning → utility scoring → shadow selection → adaptive budgeting.
pip install dspm-memory
Requires Python 3.9+. Works with any OpenAI-compatible LLM provider.
Optional: semantic ranking
pip install "dspm-memory[semantic]"
Adds sentence-transformer-based query alignment for better non-critical patch ranking (downloads PyTorch and a small embedding model on first use). Without it, DSPM falls back to a neutral alignment score — criticals are fully protected either way.
from openai import OpenAI
from dspm import DSPMMemory
# Works with OpenAI, Groq, Together, Ollama, or any OpenAI-compatible endpoint
llm = OpenAI(
api_key="sk-...",
# base_url="https://api.groq.com/openai/v1" # uncomment for Groq
)
# budget=60 so compression is visible even in this small example;
# use 200-300 for real conversations
memory = DSPMMemory(budget=60, llm_client=llm, model="gpt-4o-mini")
memory.add_turn("user", "Building a payment API. Hard rules: PCI-DSS compliant, max fee 0.5%, deadline Friday.")
memory.add_turn("assistant", "PCI-DSS needs tokenized card storage and quarterly ASV scans. Tracking the 0.5% fee cap.")
memory.add_turn("user", "Webhook timeout must be 30 seconds.")
memory.add_turn("assistant", "Webhook timeout set to 30 seconds with automatic retries.")
memory.add_turn("user", "Change the webhook timeout to 10 seconds instead — 30 is too slow.")
memory.add_turn("assistant", "Updated: webhook timeout is now 10 seconds; the 30s setting is superseded.")
context = memory.get_context(query="What are the hard requirements?")
print(context)
print(memory.stats)
Output:
[DEC] Track the 0.5% fee cap.
[CON] PCI-DSS requires tokenized card storage and ASV
[DEC] webhook timeout updated to 10 seconds
[CON] PCI-DSS compliant, max fee 0.5%, deadline Friday
{'turns': 6, 'total_patches': 4, 'critical_total': 4, 'critical_selected': 4, 'crr': 100, 'budget': 60, 'raw_tokens': 61, 'context_tokens': 56, 'trr': 9}
Note the revision: only the 10-second entry remains — the 30s setting is fully superseded. The fee cap and deadline survived compression, trr: 9 shows real token reduction even at this small scale, and crr: 100 throughout.
Save the notebook when a chat ends, load it when the next one starts — Chat 2 remembers Chat 1, and revisions supersede old values across sessions:
# Chat 1 — Monday
memory = DSPMMemory(budget=250, llm_client=llm, model="gpt-4o-mini")
memory.add_turn("user", "Building a budgeting app. Hard rules: must work offline, deadline Oct 20.")
# ... chat ...
memory.save("user_dhruv.json")
# Chat 2 — Thursday, new process, fresh start
memory = DSPMMemory(budget=250, llm_client=llm, model="gpt-4o-mini")
memory.load("user_dhruv.json")
memory.get_context(query="What were my hard rules?")
# → [CON] must work offline
# → [CON] deadline Oct 20 ← recalled from Chat 1, zero API cost
# Revisions work across chats too:
memory.add_turn("user", "The deadline moved to November 5.")
# → Oct 20 is superseded. November 5 replaces it.
Save files are portable JSON, written atomically (a crash mid-save can't corrupt the notebook), and merge-on-load means saved revisions supersede stale values. One file per user = each person's long-term memory.
[CON] and [DEC] patches are structurally protected:
30 seconds stays 30 seconds, never just 30)Ablation result: Removing the shadow-selection mechanism collapses CRR from 100% to 37.9%, isolating the guarantee to a single identifiable component.
Tested across 7 domains × 40 turns each:
| Budget | Tokens Used | TRR | CRR |
|---|---|---|---|
| 150 | 150 | 66.8% | 100% |
| 250 | 249 | 82.8% | 100% |
| 400 | 395 | 72.4% | 100% |
CRR = Critical Retention Rate. TRR = Token Reduction Ratio.
Provenance: measured on 7 hand-authored 40-turn technical dialogues (API design, ML ops, IoT, security IR, supply chain, clinical workflow, project management); extraction and judging by GLM via OpenRouter (the since-renamed "ox-alpha" model). Numbers vary by extraction model and domain — check your own with memory.stats.
# OpenAI
llm = OpenAI(api_key="sk-...")
# Groq (free tier available)
llm = OpenAI(base_url="https://api.groq.com/openai/v1", api_key="gsk-...")
# Together AI
llm = OpenAI(base_url="https://api.together.xyz/v1", api_key="...")
# Ollama (local, no key needed)
llm = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
DSPMMemory(budget, llm_client, model)| Parameter | Type | Default | Description |
|---|---|---|---|
budget | int | 250 | Maximum tokens in the compressed output |
llm_client | OpenAI | None | Any OpenAI-compatible client |
model | str | "gpt-4o-mini" | Model used for patch extraction |
The budget can be changed at any time:
memory.budget = 100 # takes effect on the next get_context() call
memory.set_budget(100) # explicit alternative — same effect
| Method | Description |
|---|---|
add_turn(role, text) | Add a conversation turn. Returns extracted patches. |
get_context(query="") | Returns compressed context string, ready for your prompt. |
set_budget(budget) | Update the token budget. Takes effect on the next get_context() call. |
save(path) | Persist the memory notebook to JSON (atomic write). |
load(path) | Load a notebook into memory (merge semantics, supersession on load). |
reset() | Clear all memory and start fresh. |
| Property | Description |
|---|---|
memory.stats | Dict with token counts, patch counts, active budget, CRR |
memory.critical_patches | List of all critical patches currently in memory |
memory.all_patches | List of every patch in memory |
DSPM: A Critical-Retention Approach to Long-Context Memory Compression for LLM Conversations
Dhruv Dubey, 2026
Zenodo: 10.5281/zenodo.19438636
| Version | Changes |
|---|---|
| 0.1.8 | Retry wording corrected — 0.1.7's wheel shipped the previous wording, caught by post-build wheel verification check immediately after upload. No code changes. |
| 0.1.7 | README-first release (0.1.6's PyPI page was frozen pre-update — built after README finalization this time). Compound units ("per minute") and key constraint nouns ("timeout", "limit") survive trimming as atomic number-spans. "now"/"moved" recognized as revision verbs, and the extractor preserves revision verbs in payloads (real-model quickstart caught a stale-value contradiction). Claims scoped to benchmarks; retry behavior accurately described. Test scripts moved to scripts/ with exposed API key revoked and stripped. semantic extra documented. Python 3.13 classifier; status → Beta. 3 new tests (27 total). |
| 0.1.6 | Cross-type revision supersession fixed: revisions sharing only 3 content words previously survived as contradictions (found in external review). Units now bind to their numbers during trimming (30 seconds, never bare 30). Automatic retry with backoff on rate limits and transient 5xx errors in add_turn(). Homepage added to PyPI metadata. 5 regression tests (24 total). |
| 0.1.5 | Budget fix: changing memory.budget was silently ignored — the engine kept an independent budget copy. Now synced on every get_context(). New set_budget() method. stats now measures the actual joined context and reports the active budget. 2 regression tests (19 total). |
| 0.1.4 | Persistence: memory.save() / memory.load() — cross-session long-term memory as portable JSON. Atomic writes, merge-on-load, revisions supersede stale values. 8 new tests (17 total). |
| 0.1.3 | Revision supersession fix: stale same-type criticals now removed when superseded. Robust normalized content-word matching. |
| 0.1.2 | Fixed critical-patch ID collisions (CRR 36% → 100% in 18-turn live test). Budget enforced on joined context string. |
| 0.1.1 | Fixed T4 dropping criticals with dependencies. Fixed T3 payload mangling. Fixed extractor schema mismatch. |
| 0.1.0 | Initial release. |
MIT © 2026 Dhruv Dubey
Training-free LLM memory compression. 100% critical retention across our benchmarks; structural guarantee under budget pressure.
Python
7
27 commits
updated Oct 5, 2026
Compress multi-turn LLM conversations by 80%+ while every constraint and decision survives.
100% critical retention across our benchmark suite — under impossibly small budgets, criticals are trimmed to their word floor and dropped only as a documented last resort.
Long conversations eat your context window. Naive truncation drops the constraint from turn 3 that the entire system depends on. DSPM fixes this.
DSPM converts each conversation turn into typed semantic patches — constraint, decision, code, entity, structure — and compresses them under a fixed token budget. Critical patches (constraints and decisions) are structurally protected: they survive compression even when everything else is trimmed.
Raw conversation (452 tokens, 18 turns)
↓
Semantic extraction → 28 patches, 18 critical
↓
7-stage compression pipeline
↓
Compressed context (249 tokens) — 100% of criticals intact
The 7-stage pipeline: dedup → slot fusion → delta encoding → causal pruning → utility scoring → shadow selection → adaptive budgeting.
pip install dspm-memory
Requires Python 3.9+. Works with any OpenAI-compatible LLM provider.
Optional: semantic ranking
pip install "dspm-memory[semantic]"
Adds sentence-transformer-based query alignment for better non-critical patch ranking (downloads PyTorch and a small embedding model on first use). Without it, DSPM falls back to a neutral alignment score — criticals are fully protected either way.
from openai import OpenAI
from dspm import DSPMMemory
# Works with OpenAI, Groq, Together, Ollama, or any OpenAI-compatible endpoint
llm = OpenAI(
api_key="sk-...",
# base_url="https://api.groq.com/openai/v1" # uncomment for Groq
)
# budget=60 so compression is visible even in this small example;
# use 200-300 for real conversations
memory = DSPMMemory(budget=60, llm_client=llm, model="gpt-4o-mini")
memory.add_turn("user", "Building a payment API. Hard rules: PCI-DSS compliant, max fee 0.5%, deadline Friday.")
memory.add_turn("assistant", "PCI-DSS needs tokenized card storage and quarterly ASV scans. Tracking the 0.5% fee cap.")
memory.add_turn("user", "Webhook timeout must be 30 seconds.")
memory.add_turn("assistant", "Webhook timeout set to 30 seconds with automatic retries.")
memory.add_turn("user", "Change the webhook timeout to 10 seconds instead — 30 is too slow.")
memory.add_turn("assistant", "Updated: webhook timeout is now 10 seconds; the 30s setting is superseded.")
context = memory.get_context(query="What are the hard requirements?")
print(context)
print(memory.stats)
Output:
[DEC] Track the 0.5% fee cap.
[CON] PCI-DSS requires tokenized card storage and ASV
[DEC] webhook timeout updated to 10 seconds
[CON] PCI-DSS compliant, max fee 0.5%, deadline Friday
{'turns': 6, 'total_patches': 4, 'critical_total': 4, 'critical_selected': 4, 'crr': 100, 'budget': 60, 'raw_tokens': 61, 'context_tokens': 56, 'trr': 9}
Note the revision: only the 10-second entry remains — the 30s setting is fully superseded. The fee cap and deadline survived compression, trr: 9 shows real token reduction even at this small scale, and crr: 100 throughout.
Save the notebook when a chat ends, load it when the next one starts — Chat 2 remembers Chat 1, and revisions supersede old values across sessions:
# Chat 1 — Monday
memory = DSPMMemory(budget=250, llm_client=llm, model="gpt-4o-mini")
memory.add_turn("user", "Building a budgeting app. Hard rules: must work offline, deadline Oct 20.")
# ... chat ...
memory.save("user_dhruv.json")
# Chat 2 — Thursday, new process, fresh start
memory = DSPMMemory(budget=250, llm_client=llm, model="gpt-4o-mini")
memory.load("user_dhruv.json")
memory.get_context(query="What were my hard rules?")
# → [CON] must work offline
# → [CON] deadline Oct 20 ← recalled from Chat 1, zero API cost
# Revisions work across chats too:
memory.add_turn("user", "The deadline moved to November 5.")
# → Oct 20 is superseded. November 5 replaces it.
Save files are portable JSON, written atomically (a crash mid-save can't corrupt the notebook), and merge-on-load means saved revisions supersede stale values. One file per user = each person's long-term memory.
[CON] and [DEC] patches are structurally protected:
30 seconds stays 30 seconds, never just 30)Ablation result: Removing the shadow-selection mechanism collapses CRR from 100% to 37.9%, isolating the guarantee to a single identifiable component.
Tested across 7 domains × 40 turns each:
| Budget | Tokens Used | TRR | CRR |
|---|---|---|---|
| 150 | 150 | 66.8% | 100% |
| 250 | 249 | 82.8% | 100% |
| 400 | 395 | 72.4% | 100% |
CRR = Critical Retention Rate. TRR = Token Reduction Ratio.
Provenance: measured on 7 hand-authored 40-turn technical dialogues (API design, ML ops, IoT, security IR, supply chain, clinical workflow, project management); extraction and judging by GLM via OpenRouter (the since-renamed "ox-alpha" model). Numbers vary by extraction model and domain — check your own with memory.stats.
# OpenAI
llm = OpenAI(api_key="sk-...")
# Groq (free tier available)
llm = OpenAI(base_url="https://api.groq.com/openai/v1", api_key="gsk-...")
# Together AI
llm = OpenAI(base_url="https://api.together.xyz/v1", api_key="...")
# Ollama (local, no key needed)
llm = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
DSPMMemory(budget, llm_client, model)| Parameter | Type | Default | Description |
|---|---|---|---|
budget | int | 250 | Maximum tokens in the compressed output |
llm_client | OpenAI | None | Any OpenAI-compatible client |
model | str | "gpt-4o-mini" | Model used for patch extraction |
The budget can be changed at any time:
memory.budget = 100 # takes effect on the next get_context() call
memory.set_budget(100) # explicit alternative — same effect
| Method | Description |
|---|---|
add_turn(role, text) | Add a conversation turn. Returns extracted patches. |
get_context(query="") | Returns compressed context string, ready for your prompt. |
set_budget(budget) | Update the token budget. Takes effect on the next get_context() call. |
save(path) | Persist the memory notebook to JSON (atomic write). |
load(path) | Load a notebook into memory (merge semantics, supersession on load). |
reset() | Clear all memory and start fresh. |
| Property | Description |
|---|---|
memory.stats | Dict with token counts, patch counts, active budget, CRR |
memory.critical_patches | List of all critical patches currently in memory |
memory.all_patches | List of every patch in memory |
DSPM: A Critical-Retention Approach to Long-Context Memory Compression for LLM Conversations
Dhruv Dubey, 2026
Zenodo: 10.5281/zenodo.19438636
| Version | Changes |
|---|---|
| 0.1.8 | Retry wording corrected — 0.1.7's wheel shipped the previous wording, caught by post-build wheel verification check immediately after upload. No code changes. |
| 0.1.7 | README-first release (0.1.6's PyPI page was frozen pre-update — built after README finalization this time). Compound units ("per minute") and key constraint nouns ("timeout", "limit") survive trimming as atomic number-spans. "now"/"moved" recognized as revision verbs, and the extractor preserves revision verbs in payloads (real-model quickstart caught a stale-value contradiction). Claims scoped to benchmarks; retry behavior accurately described. Test scripts moved to scripts/ with exposed API key revoked and stripped. semantic extra documented. Python 3.13 classifier; status → Beta. 3 new tests (27 total). |
| 0.1.6 | Cross-type revision supersession fixed: revisions sharing only 3 content words previously survived as contradictions (found in external review). Units now bind to their numbers during trimming (30 seconds, never bare 30). Automatic retry with backoff on rate limits and transient 5xx errors in add_turn(). Homepage added to PyPI metadata. 5 regression tests (24 total). |
| 0.1.5 | Budget fix: changing memory.budget was silently ignored — the engine kept an independent budget copy. Now synced on every get_context(). New set_budget() method. stats now measures the actual joined context and reports the active budget. 2 regression tests (19 total). |
| 0.1.4 | Persistence: memory.save() / memory.load() — cross-session long-term memory as portable JSON. Atomic writes, merge-on-load, revisions supersede stale values. 8 new tests (17 total). |
| 0.1.3 | Revision supersession fix: stale same-type criticals now removed when superseded. Robust normalized content-word matching. |
| 0.1.2 | Fixed critical-patch ID collisions (CRR 36% → 100% in 18-turn live test). Budget enforced on joined context string. |
| 0.1.1 | Fixed T4 dropping criticals with dependencies. Fixed T3 payload mangling. Fixed extractor schema mismatch. |
| 0.1.0 | Initial release. |
MIT © 2026 Dhruv Dubey