LUMIA-Group/MemoryDecoder-at-Scale

Python

32

1 commits

updated Jul 31, 2026

See the code

README

Memory Decoder at Scale

A Pretrained, Parametric Long-Term Memory

Rubin Wei1,2 · Jiaqi Cao1 · Jiarui Wang1,2 · Junming Zhang1 · Qipeng Guo2 · Bowen Zhou2,3 · Zhouhan Lin1,2

1LUMIA Lab, School of Artificial Intelligence, Shanghai Jiao Tong University
2Shanghai Artificial Intelligence Laboratory    3Electronic Engineering, Tsinghua University

arXiv Hugging Face Paper 🌐 Project Page Models & Data License

📖 Overview

Memory Decoder at Scale brings parametric long-term memory pretraining to language-model scale, with general memories of up to 6.9B parameters pretrained on 300B tokens. We develop a distributed Faiss indexing and retrieval pipeline, together with sparse, batch-wise loading of kNN distributions, to make memory pretraining feasible at this scale.

Across 17 benchmarks, pairing a 6.9B general memory with frozen Pythia-410M reaches an average score of 37.34, surpassing Pythia-12B while using 39% fewer total parameters. Across Qwen3 Base models from 0.6B to 14B, 1.7B domain memories improve the average over biology, law, and finance by more than 9 points at every scale.

Memory Decoder training and inference overview

🎬 Demo

A frozen backbone can be paired with swappable general or domain memories, allowing memory capacity and specialization to change without modifying the backbone.

Memory Decoder switching among general, biology, law, and finance memories

🤗 Models and Data

Released general memories, domain memories, and associated data artifacts are available in the Memory Decoder at Scale collection. Use the collection as the canonical index for public checkpoints and datasets.

🚀 Quick Start

Environment

conda env create -f environment.yml
conda activate llm-base

The environment covers preprocessing, memory training, and the core model stack. Constructing and searching a kNN datastore additionally requires a CUDA-compatible Faiss installation matching the host CUDA and Python versions.

Evaluation uses separate environments because the bundled evaluator snapshots have different dependency stacks. Create only the environment needed for the benchmarks you plan to run:

WorkflowEnvironment fileConda environment
Preprocessing and memory trainingenvironment.ymlllm-base
General, knowledge, and FinEval evaluationeval/lm-evaluation-harness/environment.ymllmeval
BioInst and LawBench evaluationeval/opencompass/environment.ymlopencompass

From the repository root, the evaluation environments can be created with:

# General, knowledge, and FinEval
conda env create -f eval/lm-evaluation-harness/environment.yml
conda activate lmeval

# BioInst and LawBench
conda env create -f eval/opencompass/environment.yml
conda activate opencompass

Activate the matching environment before running an evaluation launcher.

Evaluate a released memory

The bundled lm-evaluation-harness adapter automatically selects the task-specific interpolation weight for released Pythia memories:

cd eval/lm-evaluation-harness

lm-eval \
  --model hf-memdec \
  --model_args pretrained=EleutherAI/pythia-1.4b-deduped,memdec_path=Rubin-Wei/MemoryDecoder-Pythia-2.8B-general \
  --tasks arc_easy,piqa,mmlu \
  --batch_size 1

An explicit lmbda=... model argument remains available as an override. Evaluation environments and task-specific launchers live under eval/.

🛠️ Train a Memory Decoder

Memory pretraining has three stages: preprocess the corpus, construct retrieval-based targets, and train the parametric memory.

1. Preprocess text

For JSONL records containing a text field:

TOKENIZER_PATH=/path/to/model-or-tokenizer \
TRAIN_FILE=/path/to/train.jsonl \
TEST_FILE=/path/to/test.jsonl \
OUTPUT_DIR=/path/to/processed-domain \
bash scripts/data/preprocess_text.sh

The output is a Hugging Face DatasetDict with input_ids, attention_mask, labels, and dstore_range. The dstore_range field keeps every training example aligned with its rows in the retrieval datastore.

2. Construct kNN training targets

The pipeline saves contextual representations, builds a Faiss index, and searches the datastore to produce sparse kNN target distributions:

MODEL_PATH=/path/to/base-model \
DATASET_PATH=/path/to/processed-domain \
DSTORE_DIR=/path/to/domain-knn \
ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml \
bash scripts/data/save_knn_pipeline.sh

MODEL_FAMILY and DIMENSION are inferred from the model configuration. Common overrides include SUBSET, K, KNN_TEMP, PROBE, NCENTROIDS, CODE_SIZE, BATCH_SIZE_EVAL, and BATCH_SIZE_KNN. Set DRY_RUN=1 to inspect the generated commands without executing them.

Convert the resulting Arrow file into the memory-mapped format used during training:

INPUT_PATH=/path/to/domain-knn/knn_qwen3_train_2560.arrow \
OUTPUT_FOLDER=/path/to/domain-knn-flat \
bash scripts/data/flatten_knn_arrow.sh

Each output shard contains:

shard_000_1/
├── offset.npy
├── flatten_token_id.npy
├── flatten_prob.npy
├── label.npy
└── shape.json

[!IMPORTANT] Use the same processed dataset, tokenizer, and base model throughout target construction. Dataset order and dstore_range must remain aligned with the flattened kNN rows.

3. Train the parametric memory

MODEL_PATH=/path/to/initial-memory-model \
DATASET_PATH=/path/to/processed-domain \
KNN_PATH=/path/to/domain-knn-flat \
OUTPUT_DIR=/path/to/memory-checkpoint \
bash scripts/train/train_domain_memdec.sh

The launcher uses one GPU by default. For one-node, eight-GPU Qwen3 FSDP training, set:

ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml

Common training overrides are EPOCHS, BATCH_SIZE, GRAD_ACC_STEPS, LEARNING_RATE, WARMUP_STEPS, and ALPHA. For a multi-shard kNN store, also set NUM_SHARDS and SHARD_TEMPLATE, for example:

NUM_SHARDS=32 SHARD_TEMPLATE='shard_{:03d}_32'

📊 Evaluation

Set MODEL_PATH to evaluate a base model. Add MEMDEC_PATH to evaluate the base model together with a Memory Decoder.

Benchmark groupEntry point
General downstreambash eval/lm-evaluation-harness/scripts/general/evaluate.sh
2WikiMultihopQA, Bamboogle, HotpotQAbash eval/lm-evaluation-harness/scripts/knowledge/evaluate_multihopqa.sh
HaluEvalbash eval/lm-evaluation-harness/scripts/knowledge/evaluate_halueval.sh
FinEvalbash eval/lm-evaluation-harness/scripts/domain/evaluate_fineval.sh
BioInstbash eval/opencompass/scripts/domain/evaluate_bioinst.sh
LawBenchbash eval/opencompass/scripts/domain/evaluate_lawbench.sh

FinEval, general, and knowledge evaluations use the bundled lm-evaluation-harness tree and the lmeval environment. BioInst and LawBench use OpenCompass and the opencompass environment.

📁 Repository Structure

.
├── accelerate_config/          # Single-GPU, multi-GPU, and Qwen3 FSDP configs
├── assets/                     # README and paper figures
├── eval/
│   ├── lm-evaluation-harness/  # General, knowledge, and FinEval evaluation
│   └── opencompass/            # BioInst and LawBench evaluation
├── knn_utils/                  # Datastore, Faiss, and kNN target utilities
├── scripts/
│   ├── data/                   # Preprocessing and target construction
│   └── train/                  # Memory training launcher
├── utils/                      # Preprocessing, flattening, and loss functions
├── train_base.py               # Datastore construction entry point
└── train_memdec_trainer.py     # Sparse kNN-guided memory pretraining

🙏 Acknowledgments

This implementation builds on the excellent MemoryDecoder repository. In particular, the datastore utilities, train_base.py, and the original kNN construction pipeline were adapted from that codebase. We thank its authors for pioneering pretrained, plug-and-play parametric memory for language models.

This work is supported by the Shanghai General AI Foundation Models Program (Grant No. 2025SHZDZX025G09), with technical collaboration from the Shanghai Artificial Intelligence Laboratory.

Evaluation is built on lm-evaluation-harness and OpenCompass. We thank these communities for their open-source evaluation infrastructure.

📚 Citation

If you find Memory Decoder at Scale useful in your research, please cite:

@misc{wei2026memorydecoderscalepretrained,
      title={Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory},
      author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Junming Zhang and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
      year={2026},
      eprint={2607.27919},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2607.27919},
}

📧 Contact

Questions and discussions are welcome at weirubinn@gmail.com.

📄 License

This project is released under the Apache License 2.0.

LUMIA-Group/MemoryDecoder-at-Scale

Python

32

1 commits

updated Jul 31, 2026

See the code

README

Memory Decoder at Scale

A Pretrained, Parametric Long-Term Memory

Rubin Wei1,2 · Jiaqi Cao1 · Jiarui Wang1,2 · Junming Zhang1 · Qipeng Guo2 · Bowen Zhou2,3 · Zhouhan Lin1,2

1LUMIA Lab, School of Artificial Intelligence, Shanghai Jiao Tong University
2Shanghai Artificial Intelligence Laboratory    3Electronic Engineering, Tsinghua University

arXiv Hugging Face Paper 🌐 Project Page Models & Data License

📖 Overview

Memory Decoder at Scale brings parametric long-term memory pretraining to language-model scale, with general memories of up to 6.9B parameters pretrained on 300B tokens. We develop a distributed Faiss indexing and retrieval pipeline, together with sparse, batch-wise loading of kNN distributions, to make memory pretraining feasible at this scale.

Across 17 benchmarks, pairing a 6.9B general memory with frozen Pythia-410M reaches an average score of 37.34, surpassing Pythia-12B while using 39% fewer total parameters. Across Qwen3 Base models from 0.6B to 14B, 1.7B domain memories improve the average over biology, law, and finance by more than 9 points at every scale.

Memory Decoder training and inference overview

🎬 Demo

A frozen backbone can be paired with swappable general or domain memories, allowing memory capacity and specialization to change without modifying the backbone.

Memory Decoder switching among general, biology, law, and finance memories

🤗 Models and Data

Released general memories, domain memories, and associated data artifacts are available in the Memory Decoder at Scale collection. Use the collection as the canonical index for public checkpoints and datasets.

🚀 Quick Start

Environment

conda env create -f environment.yml
conda activate llm-base

The environment covers preprocessing, memory training, and the core model stack. Constructing and searching a kNN datastore additionally requires a CUDA-compatible Faiss installation matching the host CUDA and Python versions.

Evaluation uses separate environments because the bundled evaluator snapshots have different dependency stacks. Create only the environment needed for the benchmarks you plan to run:

WorkflowEnvironment fileConda environment
Preprocessing and memory trainingenvironment.ymlllm-base
General, knowledge, and FinEval evaluationeval/lm-evaluation-harness/environment.ymllmeval
BioInst and LawBench evaluationeval/opencompass/environment.ymlopencompass

From the repository root, the evaluation environments can be created with:

# General, knowledge, and FinEval
conda env create -f eval/lm-evaluation-harness/environment.yml
conda activate lmeval

# BioInst and LawBench
conda env create -f eval/opencompass/environment.yml
conda activate opencompass

Activate the matching environment before running an evaluation launcher.

Evaluate a released memory

The bundled lm-evaluation-harness adapter automatically selects the task-specific interpolation weight for released Pythia memories:

cd eval/lm-evaluation-harness

lm-eval \
  --model hf-memdec \
  --model_args pretrained=EleutherAI/pythia-1.4b-deduped,memdec_path=Rubin-Wei/MemoryDecoder-Pythia-2.8B-general \
  --tasks arc_easy,piqa,mmlu \
  --batch_size 1

An explicit lmbda=... model argument remains available as an override. Evaluation environments and task-specific launchers live under eval/.

🛠️ Train a Memory Decoder

Memory pretraining has three stages: preprocess the corpus, construct retrieval-based targets, and train the parametric memory.

1. Preprocess text

For JSONL records containing a text field:

TOKENIZER_PATH=/path/to/model-or-tokenizer \
TRAIN_FILE=/path/to/train.jsonl \
TEST_FILE=/path/to/test.jsonl \
OUTPUT_DIR=/path/to/processed-domain \
bash scripts/data/preprocess_text.sh

The output is a Hugging Face DatasetDict with input_ids, attention_mask, labels, and dstore_range. The dstore_range field keeps every training example aligned with its rows in the retrieval datastore.

2. Construct kNN training targets

The pipeline saves contextual representations, builds a Faiss index, and searches the datastore to produce sparse kNN target distributions:

MODEL_PATH=/path/to/base-model \
DATASET_PATH=/path/to/processed-domain \
DSTORE_DIR=/path/to/domain-knn \
ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml \
bash scripts/data/save_knn_pipeline.sh

MODEL_FAMILY and DIMENSION are inferred from the model configuration. Common overrides include SUBSET, K, KNN_TEMP, PROBE, NCENTROIDS, CODE_SIZE, BATCH_SIZE_EVAL, and BATCH_SIZE_KNN. Set DRY_RUN=1 to inspect the generated commands without executing them.

Convert the resulting Arrow file into the memory-mapped format used during training:

INPUT_PATH=/path/to/domain-knn/knn_qwen3_train_2560.arrow \
OUTPUT_FOLDER=/path/to/domain-knn-flat \
bash scripts/data/flatten_knn_arrow.sh

Each output shard contains:

shard_000_1/
├── offset.npy
├── flatten_token_id.npy
├── flatten_prob.npy
├── label.npy
└── shape.json

[!IMPORTANT] Use the same processed dataset, tokenizer, and base model throughout target construction. Dataset order and dstore_range must remain aligned with the flattened kNN rows.

3. Train the parametric memory

MODEL_PATH=/path/to/initial-memory-model \
DATASET_PATH=/path/to/processed-domain \
KNN_PATH=/path/to/domain-knn-flat \
OUTPUT_DIR=/path/to/memory-checkpoint \
bash scripts/train/train_domain_memdec.sh

The launcher uses one GPU by default. For one-node, eight-GPU Qwen3 FSDP training, set:

ACCELERATE_CONFIG=accelerate_config/qwen3_fsdp_8gpu.yaml

Common training overrides are EPOCHS, BATCH_SIZE, GRAD_ACC_STEPS, LEARNING_RATE, WARMUP_STEPS, and ALPHA. For a multi-shard kNN store, also set NUM_SHARDS and SHARD_TEMPLATE, for example:

NUM_SHARDS=32 SHARD_TEMPLATE='shard_{:03d}_32'

📊 Evaluation

Set MODEL_PATH to evaluate a base model. Add MEMDEC_PATH to evaluate the base model together with a Memory Decoder.

Benchmark groupEntry point
General downstreambash eval/lm-evaluation-harness/scripts/general/evaluate.sh
2WikiMultihopQA, Bamboogle, HotpotQAbash eval/lm-evaluation-harness/scripts/knowledge/evaluate_multihopqa.sh
HaluEvalbash eval/lm-evaluation-harness/scripts/knowledge/evaluate_halueval.sh
FinEvalbash eval/lm-evaluation-harness/scripts/domain/evaluate_fineval.sh
BioInstbash eval/opencompass/scripts/domain/evaluate_bioinst.sh
LawBenchbash eval/opencompass/scripts/domain/evaluate_lawbench.sh

FinEval, general, and knowledge evaluations use the bundled lm-evaluation-harness tree and the lmeval environment. BioInst and LawBench use OpenCompass and the opencompass environment.

📁 Repository Structure

.
├── accelerate_config/          # Single-GPU, multi-GPU, and Qwen3 FSDP configs
├── assets/                     # README and paper figures
├── eval/
│   ├── lm-evaluation-harness/  # General, knowledge, and FinEval evaluation
│   └── opencompass/            # BioInst and LawBench evaluation
├── knn_utils/                  # Datastore, Faiss, and kNN target utilities
├── scripts/
│   ├── data/                   # Preprocessing and target construction
│   └── train/                  # Memory training launcher
├── utils/                      # Preprocessing, flattening, and loss functions
├── train_base.py               # Datastore construction entry point
└── train_memdec_trainer.py     # Sparse kNN-guided memory pretraining

🙏 Acknowledgments

This implementation builds on the excellent MemoryDecoder repository. In particular, the datastore utilities, train_base.py, and the original kNN construction pipeline were adapted from that codebase. We thank its authors for pioneering pretrained, plug-and-play parametric memory for language models.

This work is supported by the Shanghai General AI Foundation Models Program (Grant No. 2025SHZDZX025G09), with technical collaboration from the Shanghai Artificial Intelligence Laboratory.

Evaluation is built on lm-evaluation-harness and OpenCompass. We thank these communities for their open-source evaluation infrastructure.

📚 Citation

If you find Memory Decoder at Scale useful in your research, please cite:

@misc{wei2026memorydecoderscalepretrained,
      title={Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory},
      author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Junming Zhang and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
      year={2026},
      eprint={2607.27919},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2607.27919},
}

📧 Contact

Questions and discussions are welcome at weirubinn@gmail.com.

📄 License

This project is released under the Apache License 2.0.

Languages

Python

99.2%