A Pretrained, Parametric Long-Term Memory
Rubin Wei1,2 · Jiaqi Cao1 · Jiarui Wang1,2 · Junming Zhang1 · Qipeng Guo2 · Bowen Zhou2,3 · Zhouhan Lin1,2
1LUMIA Lab, School of Artificial Intelligence, Shanghai Jiao Tong University
2Shanghai Artificial Intelligence Laboratory
3Electronic Engineering, Tsinghua University
Memory Decoder at Scale brings parametric long-term memory pretraining to language-model scale, with general memories of up to 6.9B parameters pretrained on 300B tokens. We develop a distributed Faiss indexing and retrieval pipeline, together with sparse, batch-wise loading of kNN distributions, to make memory pretraining feasible at this scale.
Across 17 benchmarks, pairing a 6.9B general memory with frozen Pythia-410M reaches an average score of 37.34, surpassing Pythia-12B while using 39% fewer total parameters. Across Qwen3 Base models from 0.6B to 14B, 1.7B domain memories improve the average over biology, law, and finance by more than 9 points at every scale.
A frozen backbone can be paired with swappable general or domain memories, allowing memory capacity and specialization to change without modifying the backbone.
Released general memories, domain memories, and associated data artifacts are available in the Memory Decoder at Scale collection. Use the collection as the canonical index for public checkpoints and datasets.
conda env create -f environment.yml
conda activate llm-base
The environment covers preprocessing, memory training, and the core model stack. Constructing and searching a kNN datastore additionally requires a CUDA-compatible Faiss installation matching the host CUDA and Python versions.
Evaluation uses separate environments because the bundled evaluator snapshots have different dependency stacks. Create only the environment needed for the benchmarks you plan to run:
| Workflow | Environment file | Conda environment |
|---|---|---|
| Preprocessing and memory training | environment.yml | llm-base |
| General, knowledge, and FinEval evaluation | eval/lm-evaluation-harness/environment.yml | lmeval |
| BioInst and LawBench evaluation | eval/opencompass/environment.yml | opencompass |
From the repository root, the evaluation environments can be created with:
# General, knowledge, and FinEval
conda env create -f eval/lm-evaluation-harness/environment.yml
conda activate lmeval
# BioInst and LawBench
conda env create -f eval/opencompass/environment.yml
conda activate opencompass
Activate the matching environment before running an evaluation launcher.
The bundled lm-evaluation-harness adapter automatically selects the task-specific interpolation weight for released Pythia memories:
cd eval/lm-evaluation-harness
lm-eval \
--model hf-memdec \
--model_args pretrained=EleutherAI/pythia-1.4b-deduped,memdec_path=Rubin-Wei/MemoryDecoder-Pythia-2.8B-general \
--tasks arc_easy,piqa,mmlu \
--batch_size 1
An explicit lmbda=... model argument remains available as an override.
Evaluation environments and task-specific launchers live under eval/.
Memory pretraining has three stages: preprocess the corpus, construct retrieval-based targets, and train the parametric memory.
For JSONL records containing a text field:
TOKENIZER_PATH=/path/to/model-or-tokenizer \
TRAIN_FILE=/path/to/train.jsonl \
TEST_FILE=/path/to/test.jsonl \
OUTPUT_DIR=/path/to/processed-domain \
bash scripts/data/preprocess_text.sh
The output is a Hugging Face DatasetDict with input_ids,
attention_mask, labels, and dstore_range. The dstore_range field keeps
every training example aligned with its rows in the retrieval datastore.
The pipeline saves contextual representations, builds a Faiss index, and searches the datastore to produce sparse kNN target distributions:
MODEL_PATH=/path/to/base-model \
DATASET_PATH=/path/to/processed-domain \
DSTORE_DIR=/path/to/domain-knn \
ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml \
bash scripts/data/save_knn_pipeline.sh
MODEL_FAMILY and DIMENSION are inferred from the model configuration.
Common overrides include SUBSET, K, KNN_TEMP, PROBE, NCENTROIDS,
CODE_SIZE, BATCH_SIZE_EVAL, and BATCH_SIZE_KNN. Set DRY_RUN=1 to
inspect the generated commands without executing them.
Convert the resulting Arrow file into the memory-mapped format used during training:
INPUT_PATH=/path/to/domain-knn/knn_qwen3_train_2560.arrow \
OUTPUT_FOLDER=/path/to/domain-knn-flat \
bash scripts/data/flatten_knn_arrow.sh
Each output shard contains:
shard_000_1/
├── offset.npy
├── flatten_token_id.npy
├── flatten_prob.npy
├── label.npy
└── shape.json
[!IMPORTANT] Use the same processed dataset, tokenizer, and base model throughout target construction. Dataset order and
dstore_rangemust remain aligned with the flattened kNN rows.
MODEL_PATH=/path/to/initial-memory-model \
DATASET_PATH=/path/to/processed-domain \
KNN_PATH=/path/to/domain-knn-flat \
OUTPUT_DIR=/path/to/memory-checkpoint \
bash scripts/train/train_domain_memdec.sh
The launcher uses one GPU by default. For one-node, eight-GPU Qwen3 FSDP training, set:
ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml
Common training overrides are EPOCHS, BATCH_SIZE, GRAD_ACC_STEPS,
LEARNING_RATE, WARMUP_STEPS, and ALPHA. For a multi-shard kNN store,
also set NUM_SHARDS and SHARD_TEMPLATE, for example:
NUM_SHARDS=32 SHARD_TEMPLATE='shard_{:03d}_32'
Set MODEL_PATH to evaluate a base model. Add MEMDEC_PATH to evaluate the
base model together with a Memory Decoder.
| Benchmark group | Entry point |
|---|---|
| General downstream | bash eval/lm-evaluation-harness/scripts/general/evaluate.sh |
| 2WikiMultihopQA, Bamboogle, HotpotQA | bash eval/lm-evaluation-harness/scripts/knowledge/evaluate_multihopqa.sh |
| HaluEval | bash eval/lm-evaluation-harness/scripts/knowledge/evaluate_halueval.sh |
| FinEval | bash eval/lm-evaluation-harness/scripts/domain/evaluate_fineval.sh |
| BioInst | bash eval/opencompass/scripts/domain/evaluate_bioinst.sh |
| LawBench | bash eval/opencompass/scripts/domain/evaluate_lawbench.sh |
FinEval, general, and knowledge evaluations use the bundled
lm-evaluation-harness tree and the lmeval environment. BioInst and LawBench
use OpenCompass and the opencompass environment.
.
├── accelerate_config/ # Single-GPU, multi-GPU, and Qwen3 FSDP configs
├── assets/ # README and paper figures
├── eval/
│ ├── lm-evaluation-harness/ # General, knowledge, and FinEval evaluation
│ └── opencompass/ # BioInst and LawBench evaluation
├── knn_utils/ # Datastore, Faiss, and kNN target utilities
├── scripts/
│ ├── data/ # Preprocessing and target construction
│ └── train/ # Memory training launcher
├── utils/ # Preprocessing, flattening, and loss functions
├── train_base.py # Datastore construction entry point
└── train_memdec_trainer.py # Sparse kNN-guided memory pretraining
This implementation builds on the excellent
MemoryDecoder repository.
In particular, the datastore utilities, train_base.py, and the original kNN
construction pipeline were adapted from that codebase. We thank its authors
for pioneering pretrained, plug-and-play parametric memory for language
models.
This work is supported by the Shanghai General AI Foundation Models Program (Grant No. 2025SHZDZX025G09), with technical collaboration from the Shanghai Artificial Intelligence Laboratory.
Evaluation is built on lm-evaluation-harness and OpenCompass. We thank these communities for their open-source evaluation infrastructure.
If you find Memory Decoder at Scale useful in your research, please cite:
@misc{wei2026memorydecoderscalepretrained,
title={Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory},
author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Junming Zhang and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
year={2026},
eprint={2607.27919},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.27919},
}
Questions and discussions are welcome at weirubinn@gmail.com.
This project is released under the Apache License 2.0.
Python
99.2%
A Pretrained, Parametric Long-Term Memory
Rubin Wei1,2 · Jiaqi Cao1 · Jiarui Wang1,2 · Junming Zhang1 · Qipeng Guo2 · Bowen Zhou2,3 · Zhouhan Lin1,2
1LUMIA Lab, School of Artificial Intelligence, Shanghai Jiao Tong University
2Shanghai Artificial Intelligence Laboratory
3Electronic Engineering, Tsinghua University
Memory Decoder at Scale brings parametric long-term memory pretraining to language-model scale, with general memories of up to 6.9B parameters pretrained on 300B tokens. We develop a distributed Faiss indexing and retrieval pipeline, together with sparse, batch-wise loading of kNN distributions, to make memory pretraining feasible at this scale.
Across 17 benchmarks, pairing a 6.9B general memory with frozen Pythia-410M reaches an average score of 37.34, surpassing Pythia-12B while using 39% fewer total parameters. Across Qwen3 Base models from 0.6B to 14B, 1.7B domain memories improve the average over biology, law, and finance by more than 9 points at every scale.
A frozen backbone can be paired with swappable general or domain memories, allowing memory capacity and specialization to change without modifying the backbone.
Released general memories, domain memories, and associated data artifacts are available in the Memory Decoder at Scale collection. Use the collection as the canonical index for public checkpoints and datasets.
conda env create -f environment.yml
conda activate llm-base
The environment covers preprocessing, memory training, and the core model stack. Constructing and searching a kNN datastore additionally requires a CUDA-compatible Faiss installation matching the host CUDA and Python versions.
Evaluation uses separate environments because the bundled evaluator snapshots have different dependency stacks. Create only the environment needed for the benchmarks you plan to run:
| Workflow | Environment file | Conda environment |
|---|---|---|
| Preprocessing and memory training | environment.yml | llm-base |
| General, knowledge, and FinEval evaluation | eval/lm-evaluation-harness/environment.yml | lmeval |
| BioInst and LawBench evaluation | eval/opencompass/environment.yml | opencompass |
From the repository root, the evaluation environments can be created with:
# General, knowledge, and FinEval
conda env create -f eval/lm-evaluation-harness/environment.yml
conda activate lmeval
# BioInst and LawBench
conda env create -f eval/opencompass/environment.yml
conda activate opencompass
Activate the matching environment before running an evaluation launcher.
The bundled lm-evaluation-harness adapter automatically selects the task-specific interpolation weight for released Pythia memories:
cd eval/lm-evaluation-harness
lm-eval \
--model hf-memdec \
--model_args pretrained=EleutherAI/pythia-1.4b-deduped,memdec_path=Rubin-Wei/MemoryDecoder-Pythia-2.8B-general \
--tasks arc_easy,piqa,mmlu \
--batch_size 1
An explicit lmbda=... model argument remains available as an override.
Evaluation environments and task-specific launchers live under eval/.
Memory pretraining has three stages: preprocess the corpus, construct retrieval-based targets, and train the parametric memory.
For JSONL records containing a text field:
TOKENIZER_PATH=/path/to/model-or-tokenizer \
TRAIN_FILE=/path/to/train.jsonl \
TEST_FILE=/path/to/test.jsonl \
OUTPUT_DIR=/path/to/processed-domain \
bash scripts/data/preprocess_text.sh
The output is a Hugging Face DatasetDict with input_ids,
attention_mask, labels, and dstore_range. The dstore_range field keeps
every training example aligned with its rows in the retrieval datastore.
The pipeline saves contextual representations, builds a Faiss index, and searches the datastore to produce sparse kNN target distributions:
MODEL_PATH=/path/to/base-model \
DATASET_PATH=/path/to/processed-domain \
DSTORE_DIR=/path/to/domain-knn \
ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml \
bash scripts/data/save_knn_pipeline.sh
MODEL_FAMILY and DIMENSION are inferred from the model configuration.
Common overrides include SUBSET, K, KNN_TEMP, PROBE, NCENTROIDS,
CODE_SIZE, BATCH_SIZE_EVAL, and BATCH_SIZE_KNN. Set DRY_RUN=1 to
inspect the generated commands without executing them.
Convert the resulting Arrow file into the memory-mapped format used during training:
INPUT_PATH=/path/to/domain-knn/knn_qwen3_train_2560.arrow \
OUTPUT_FOLDER=/path/to/domain-knn-flat \
bash scripts/data/flatten_knn_arrow.sh
Each output shard contains:
shard_000_1/
├── offset.npy
├── flatten_token_id.npy
├── flatten_prob.npy
├── label.npy
└── shape.json
[!IMPORTANT] Use the same processed dataset, tokenizer, and base model throughout target construction. Dataset order and
dstore_rangemust remain aligned with the flattened kNN rows.
MODEL_PATH=/path/to/initial-memory-model \
DATASET_PATH=/path/to/processed-domain \
KNN_PATH=/path/to/domain-knn-flat \
OUTPUT_DIR=/path/to/memory-checkpoint \
bash scripts/train/train_domain_memdec.sh
The launcher uses one GPU by default. For one-node, eight-GPU Qwen3 FSDP training, set:
ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml
Common training overrides are EPOCHS, BATCH_SIZE, GRAD_ACC_STEPS,
LEARNING_RATE, WARMUP_STEPS, and ALPHA. For a multi-shard kNN store,
also set NUM_SHARDS and SHARD_TEMPLATE, for example:
NUM_SHARDS=32 SHARD_TEMPLATE='shard_{:03d}_32'
Set MODEL_PATH to evaluate a base model. Add MEMDEC_PATH to evaluate the
base model together with a Memory Decoder.
| Benchmark group | Entry point |
|---|---|
| General downstream | bash eval/lm-evaluation-harness/scripts/general/evaluate.sh |
| 2WikiMultihopQA, Bamboogle, HotpotQA | bash eval/lm-evaluation-harness/scripts/knowledge/evaluate_multihopqa.sh |
| HaluEval | bash eval/lm-evaluation-harness/scripts/knowledge/evaluate_halueval.sh |
| FinEval | bash eval/lm-evaluation-harness/scripts/domain/evaluate_fineval.sh |
| BioInst | bash eval/opencompass/scripts/domain/evaluate_bioinst.sh |
| LawBench | bash eval/opencompass/scripts/domain/evaluate_lawbench.sh |
FinEval, general, and knowledge evaluations use the bundled
lm-evaluation-harness tree and the lmeval environment. BioInst and LawBench
use OpenCompass and the opencompass environment.
.
├── accelerate_config/ # Single-GPU, multi-GPU, and Qwen3 FSDP configs
├── assets/ # README and paper figures
├── eval/
│ ├── lm-evaluation-harness/ # General, knowledge, and FinEval evaluation
│ └── opencompass/ # BioInst and LawBench evaluation
├── knn_utils/ # Datastore, Faiss, and kNN target utilities
├── scripts/
│ ├── data/ # Preprocessing and target construction
│ └── train/ # Memory training launcher
├── utils/ # Preprocessing, flattening, and loss functions
├── train_base.py # Datastore construction entry point
└── train_memdec_trainer.py # Sparse kNN-guided memory pretraining
This implementation builds on the excellent
MemoryDecoder repository.
In particular, the datastore utilities, train_base.py, and the original kNN
construction pipeline were adapted from that codebase. We thank its authors
for pioneering pretrained, plug-and-play parametric memory for language
models.
This work is supported by the Shanghai General AI Foundation Models Program (Grant No. 2025SHZDZX025G09), with technical collaboration from the Shanghai Artificial Intelligence Laboratory.
Evaluation is built on lm-evaluation-harness and OpenCompass. We thank these communities for their open-source evaluation infrastructure.
If you find Memory Decoder at Scale useful in your research, please cite:
@misc{wei2026memorydecoderscalepretrained,
title={Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory},
author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Junming Zhang and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
year={2026},
eprint={2607.27919},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.27919},
}
Questions and discussions are welcome at weirubinn@gmail.com.
This project is released under the Apache License 2.0.
Python
99.2%