Awesome papers on personalization in LLMs/MLLMs: memory, alignment, retrieval, and evaluation.
See the codeThe most comprehensive survey and frontier tracking repository for personalized LLMs and MLLMs, covering personalized memory, alignment, retrieval, and evaluation.
Personalized LLMs and MLLMs aim to move beyond one-size-fits-all assistants. Instead of only optimizing for average human preference, they need to model a specific user: long-term goals, evolving preferences, implicit personas, multimodal context, and when personalization should or should not be applied.
This repository tracks papers, benchmarks, datasets, and systems around four connected research directions:
| Direction | Core Question |
|---|---|
| Personalized Memory | What should an agent store, update, retrieve, compress, and forget? |
| Personalized Alignment | How can a model adapt to individual preferences, personalities, and contexts? |
| Personalized Retrieval | How should systems select the right user context, memory, and evidence? |
| Personalized Evaluation | How do we evaluate long-term, dynamic, implicit, and multimodal personalization? |
What should an agent store, update, retrieve, compress, and forget?
How can a model adapt to individual preferences, personalities, and contexts?
How should systems select the right user context, memory, and evidence?
| Date | Paper Title | Venue | Publication | GitHub / Stars |
|---|---|---|---|---|
| 2026.05 | OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation | University of Illinois Urbana-Champaign / Meta | arXiv | - |
| 2026.04 | MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models | Carnegie Mellon University / Google DeepMind / TIGER-Lab / University of Waterloo | arXiv | MMEB Leaderboard |
| 2025.10 | Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video | NVIDIA | arXiv | nvidia/omni-embed-nemotron-3b |
| 2025.10 | SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model | ByteDance Douyin Content | arXiv | BytedanceDouyinContent/SAIL-Embedding |
| 2025.09 | WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM | TCL Research America | ICLR 2026 | TCL606/WAVE |
| 2025.01 | VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks | TIGER-Lab / University of Waterloo | ICLR 2025 | TIGER-AI-Lab/VLM2Vec |
How do we evaluate long-term, dynamic, implicit, and multimodal personalization?
If we missed any relevant work, please feel free to open an issue to contact us.
If this list is useful, please consider citing or starring the repository after publication.
53 commits
Awesome papers on personalization in LLMs/MLLMs: memory, alignment, retrieval, and evaluation.
See the codeThe most comprehensive survey and frontier tracking repository for personalized LLMs and MLLMs, covering personalized memory, alignment, retrieval, and evaluation.
Personalized LLMs and MLLMs aim to move beyond one-size-fits-all assistants. Instead of only optimizing for average human preference, they need to model a specific user: long-term goals, evolving preferences, implicit personas, multimodal context, and when personalization should or should not be applied.
This repository tracks papers, benchmarks, datasets, and systems around four connected research directions:
| Direction | Core Question |
|---|---|
| Personalized Memory | What should an agent store, update, retrieve, compress, and forget? |
| Personalized Alignment | How can a model adapt to individual preferences, personalities, and contexts? |
| Personalized Retrieval | How should systems select the right user context, memory, and evidence? |
| Personalized Evaluation | How do we evaluate long-term, dynamic, implicit, and multimodal personalization? |
What should an agent store, update, retrieve, compress, and forget?
How can a model adapt to individual preferences, personalities, and contexts?
How should systems select the right user context, memory, and evidence?
| Date | Paper Title | Venue | Publication | GitHub / Stars |
|---|---|---|---|---|
| 2026.05 | OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation | University of Illinois Urbana-Champaign / Meta | arXiv | - |
| 2026.04 | MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models | Carnegie Mellon University / Google DeepMind / TIGER-Lab / University of Waterloo | arXiv | MMEB Leaderboard |
| 2025.10 | Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video | NVIDIA | arXiv | nvidia/omni-embed-nemotron-3b |
| 2025.10 | SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model | ByteDance Douyin Content | arXiv | BytedanceDouyinContent/SAIL-Embedding |
| 2025.09 | WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM | TCL Research America | ICLR 2026 | TCL606/WAVE |
| 2025.01 | VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks | TIGER-Lab / University of Waterloo | ICLR 2025 | TIGER-AI-Lab/VLM2Vec |
How do we evaluate long-term, dynamic, implicit, and multimodal personalization?
If we missed any relevant work, please feel free to open an issue to contact us.
If this list is useful, please consider citing or starring the repository after publication.
53 commits