
This repository contains a collection of papers and resources on Personalized Content Synthesis (PCS) with Diffusion Model.
If you find the information in our paper useful for your research, please consider citing it in your work. Thank you!
@misc{zhang2024survey,
title={A Survey on Personalized Content Synthesis with Diffusion Models},
author={Xulu Zhang and Xiao-Yong Wei and Wengyu Zhang and Jinlin Wu and Zhaoxiang Zhang and Zhen Lei and Qing Li},
year={2024},
eprint={2405.05538},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
To uniformly evaluate Personalized Content Synthesis (PCS) tasks, we introduces a comprehensive evaluation dataset designed for the most common personalized generation tasks, object and face personalization.
Download Link: PCS-dataset
We evaluate existing representative PCS methods based our unified test dataset. The evaluation results and settings are shown in the below table. For more details, please see our survey paper.
| Type | Methods | Framework | Backbone | CLIP-T | CLIP-I |
|---|---|---|---|---|---|
| Object | Textual Inversion | TTF | SD 1.5 | 0.199 | 0.749 |
| Dreambooth | TTF | SD 1.5 | 0.286 | 0.772 | |
| P+ | TTF | SD 1.4 | 0.244 | 0.643 | |
| Custom Diffusion | TTF | SD 1.4 | 0.307 | 0.722 | |
| NeTI | TTF | SD 1.4 | 0.283 | 0.801 | |
| SVDiff | TTF | SD 1.5 | 0.282 | 0.776 | |
| Perfusion | TTF | SD 1.5 | 0.273 | 0.691 | |
| ELITE | PTA | SD 1.4 | 0.292 | 0.765 | |
| BLIP-Diffusion | PTA | SD 1.5 | 0.292 | 0.772 | |
| IP-Adapter | PTA | SD 1.5 | 0.272 | 0.825 | |
| SSR Encoder | PTA | SD 1.5 | 0.288 | 0.792 | |
| MoMA | PTA | SD 1.5 | 0.322 | 0.748 | |
| Diptych Prompting | PTA | FLUX 1.0 dev | 0.327 | 0.722 | |
| ฮป-eclipse | PTA | Kandinsky 2.2 | 0.272 | 0.824 | |
| MS-Diffusion | PTA | SDXL | 0.298 | 0.777 | |
| Face | CrossInitialization | TTF | SD 2.1 | 0.261 | 0.469 |
| Face2Diffusion | PTA | SD 1.4 | 0.265 | 0.588 | |
| SSR Encoder | PTA | SD 1.5 | 0.233 | 0.490 | |
| FastComposer | PTA | SD 1.5 | 0.230 | 0.516 | |
| IP-Adapter | PTA | SD 1.5 | 0.292 | 0.462 | |
| IP-Adapter | PTA | SDXL | 0.292 | 0.642 | |
| PhotoMaker | PTA | SDXL | 0.311 | 0.547 | |
| InstantID | PTA | SDXL | 0.278 | 0.707 |
TTF: Test-time Fine-tuning, PTA: Pre-trained Adaptation
| Title | Venue | Date | Links |
|---|---|---|---|
| An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion | ICLR 2023 | 2022 08-02 | Code Paper |
| DreamBooth: Fine-Tuning Text-to-Image Diffusion Models for Subject-Driven Generation | CVPR 2023 | 2022 08-25 | Code Paper |
| Re-Imagen: Retrieval-Augmented Text-to-Image Generator | ICLR 2023 | 2022 09-29 | โ Paper |
| Versatile Diffusion: Text, Images, and Variations All in One Diffusion Model | ICCV 2023 | 2022 11-15 | Code Paper |
| DreamArtist: Towards Controllable One-Shot Text-to-Image Generation via Positive-Negative Prompt-Tuning | arXiv 2022 | 2022 11-21 | Code Paper |
| Is This Loss Informative? Faster Text-to-Image Customization by Tracking Objective Dynamics | NeurIPS 2023 | 2023 02-09 | Code Paper |
| Encoder-Based Domain Tuning for Fast Personalization of Text-to-Image Models | ACM Trans on Graphics | 2023 02-23 | Code Paper |
| ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation | ICCV 2023 | 2023 02-27 | Code Paper |
| Highly Personalized Text Embedding for Image Manipulation by Stable Diffusion | arXiv | 2023 03-15 | Code Paper |
| Unified Multi-Modal Latent Diffusion for Joint Subject and Text Conditional Image Generation | arXiv | 2023 03-16 | โ Paper |
| P+: Extended Textual Conditioning in Text-to-Image Generation | arXiv | 2023 03-16 | Code Paper |
| A Closer Look at Parameter-Efficient Tuning in Diffusion Models | arXiv | 2023 03-31 | Code Paper |
| Subject-Driven Text-to-Image Generation via Apprenticeship Learning | NIPS 2023 | 2023 04-01 | โ Paper |
| Taming Encoder for Zero Fine-Tuning Image Customization with Text-to-Image Diffusion Models | arXiv | 2023 04-05 | โ Paper |
| InstantBooth: Personalized Text-to-Image Generation Without Test-Time Finetuning | arXiv | 2023 04-06 | Code Paper |
| Controllable Textual Inversion for Personalized Text-to-Image Generation | arXiv | 2023 04-11 | Code Paper |
| Gradient-Free Textual Inversion | ACM MM 2023 | 2023 04-12 | Code Paper |
| Personalize Segment Anything Model with One Shot | ICLR 2024 | 2023 05-04 | Code Paper |
| DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation | ICLR 2024 | 2023 05-05 | Code Paper |
| BLIP-Diffusion: Pre-Trained Subject Representation for Controllable Text-to-Image Generation and Editing | NIPS 2023 | 2023 05-24 | Code Paper |
| A Neural Space-Time Representation for Text-to-Image Personalization | SIGGRAPH Asia 2023 | 2023 05-24 | Code Paper |
| Prospect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models | ACM Trans on Graphics | 2023 05-25 | Code Paper |
| Break-a-Scene: Extracting Multiple Concepts from a Single Image | SIGGRAPH ASIA 2023 | 2023 05-25 | Code Paper |
| COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models | arXiv | 2023 05-26 | Code Paper |
| ViCo: Plug-and-Play Visual Condition for Personalized Text-to-Image Generation | arXiv | 2023 06-01 | Code Paper |
| Controlling Text-to-Image Diffusion by Orthogonal Fine-Tuning | arXiv | 2023 06-12 | Code Paper |
| Domain-Agnostic Tuning-Encoder for Fast Personalization of Text-to-Image Models | SIGGRAPH 2023 | 2023 07-13 | Code Paper |
| IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models | arXiv | 2023 08-13 | Code Paper |
| Navigating Text-to-Image Customization: From Lycoris Fine-Tuning to Model Evaluation | ICLR 2024 | 2023 09-26 | Code Paper |
| Kosmos-G: Generating Images in Context with Multimodal Large Language Models | arXiv | 2023 10-04 | Code Paper |
| Personalized Text-to-Image Model Enhancement Strategies: SOD Preprocessing and CNN Local Feature Integration | arXiv | 2023 10-26 | โ Paper |
| A Data Perspective on Enhanced Identity Preservation for Diffusion Personalization | ICLR 2024 | 2023 11-07 | Code Paper |
| DIFFNAT: Improving Diffusion Image Quality Using Natural Image Statistics | arXiv | 2023 11-16 | โ Paper |
| An Image is Worth Multiple Words: Multi-attribute Inversion for Constrained Text-to-Image Synthesis | arXiv | 2023 11-20 | โ Paper |
| LEGO: Learning to Disentangle and Invert Concepts Beyond Object Appearance in Text-to-Image Diffusion Models | arXiv | 2023 11-23 | Code Paper |
| Catversion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization | arXiv | 2023 11-24 | Code Paper |
| CLiC: Concept Learning in Context | CVPR 2024 | 2023 11-28 | Code Paper |
| HiFi Tuner: High-Fidelity Subject-Driven Fine-Tuning for Diffusion Models | arXiv | 2023 11-30 | โ Paper |
| InstructBooth: Instruction-Following Personalized Text-to-Image Generation | arXiv | 2023 12-04 | โ Paper |
| Customization Assistant for Text-to-Image Generation | CVPR 2024 | 2023 12-05 | โ Paper |
| Decoupled Textual Embeddings for Customized Image Generation | AAAI 2024 | 2023 12-19 | Code Paper |
| Towards Accurate Guided Diffusion Sampling through Symplectic Adjoint Method | arXiv | 2023 12-19 | Code Paper |
| DreamDistribution: Prompt Distribution Learning for Text-to-Image Diffusion Models | arXiv | 2023 12-21 | Code Paper |
| DreamTuner: Single Image is Enough for Subject-Driven Generation | arXiv | 2023 12-21 | Code Paper |
| BootPIG: Bootstrapping Zero-Shot Personalized Image Generation Capabilities in Pretrained Diffusion Models | arXiv | 2024 01-25 | โ Paper |
| Object-Driven One-Shot Fine-Tuning of Text-to-Image Diffusion with Prototypical Embedding | arXiv | 2024 01-28 | โ Paper |
| DisenDreamer: Subject-Driven Text-to-Image Generation with Sample-aware Disentangled Tuning | arXiv | 2024 02-26 | โ Paper |
| Infusion: Preventing Customized Text-to-Image Diffusion from Overfitting | arXiv | 2024 04-22 | โ Paper |
| Customizing Text-to-Image Models with a Single Image Pair | arXiv | 2024 05-02 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Multi-Concept Customization of Text-to-Image Diffusion | CVPR 2023 | 2022 12-08 | Code Paper |
| CONES: Concept Neurons in Diffusion Models for Customized Generation | ICML 2023 | 2023 03-09 | Code Paper |
| SVDiff: Compact Parameter Space for Diffusion Fine-Tuning | ICCV 2023 | 2023 03-20 | Code Paper |
| Key-Locked Rank One Editing for Text-to-Image Personalization | SIGGRAPH 2023 | 2023 05-02 | โ Paper |
| Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models | NIPS 2023 | 2023 05-29 | Code Paper |
| CONES 2: Customizable Image Synthesis with Multiple Subjects | NIPS 2023 | 2023 05-30 | Code Paper |
| Generate Anything Anywhere in Any Scene | arXiv | 2023 06-29 | Code Paper |
| AnyDoor: Zero-Shot Object-Level Image Customization | arXiv | 2023 07-18 | Code Paper |
| Subject-Diffusion: Open Domain Personalized Text-to-Image Generation Without Test-Time Fine-Tuning | arXiv | 2023 07-21 | Code Paper |
| CustomNet: Zero-Shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models | arXiv | 2023 10-30 | Code Paper |
| Compositional Inversion for Stable Diffusion Models | AAAI 2024 | 2023 12-13 | Code Paper |
| Visual Concept-Driven Image Generation with Text-to-Image Diffusion Model | arXiv | 2024 02-18 | โ Paper |
| MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis | arXiv | 2024 02-27 | โ Paper |
| Multi-Object Editing in Personalized Text-To-Image Diffusion Model Via Segmentation Guidance | arXiv | 2024 03-18 | โ Paper |
| MC2: Multi-concept Guidance for Customized Multi-concept Generation | arXiv | 2024 04-12 | โ Paper |
| MultiBooth: Towards Generating All Your Concepts in an Image from Text | arXiv | 2024 04-22 | โ Paper |
| MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models | arXiv | 2024 10-06 | Code Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| StyleDrop: Text-to-Image Synthesis of Any Style | NIPS 2023 | 2023 06-01 | Code Paper |
| StyleAdapter: A Single-Pass LoRA-Free Model for Stylized Image Generation | ICLR 2024 | 2023 09-04 | โ Paper |
| StyleBoost: A Study of Personalizing Text-to-Image Generation in Any Style using DreamBooth | ICTC 2023 | 2023 10-13 | โ Paper |
| ArtAdapter: Text-to-Image Style Transfer Using Multi-Level Style Encoder and Explicit Adaptation | arXiv | 2023 12-04 | Code Paper |
| Style Aligned Image Generation via Shared Attention | CVPR 2024 | 2023 12-04 | Code Paper |
| Generative Active Learning for Image Synthesis Personalization | arXiv | 2024 03-22 | Code Paper |
| Text-to-Image Synthesis for Any Artistic Styles: Advancements in Personalized Artistic Image Generation via Subdivision and Dual Binding | arXiv | 2024 04-08 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Identity Encoder for Personalized Diffusion | CoRR 2023 | 2023 04-14 | โ Paper |
| FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention | CoRR 2023 | 2023 05-21 | Code Paper |
| Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach | arXiv | 2023 05-23 | โ Paper |
| Inserting Anybody in Diffusion Models via Celeb Basis | NIPS 2023 | 2023 06-01 | Code Paper |
| Face0: Instantaneously Conditioning a Text-to-Image Model on a Face | SIGGRAPH 2023 | 2023 06-11 | โ Paper |
| DreamIdentity: Improved Editability for Efficient Face-Identity Preserved Image Generation | arXiv | 2023 07-01 | Code Paper |
| HyperDreamBooth: Hypernetworks for Fast Personalization of Text-to-Image Models | arXiv | 2023 07-13 | Code Paper |
| Identity-Preserving Aging of Face Images via Latent Diffusion Models | IJCB 2023 | 2023 07-17 | Code Paper |
| Magicapture: High-Resolution Multi-Concept Portrait Customization | arXiv | 2023 09-13 | Code Paper |
| High-Fidelity Person-Centric Subject-to-Image Synthesis | CVPR 2024 | 2023 11-17 | โ Paper |
| When StyleGAN Meets Stable Diffusion: A W+ Adapter for Personalized Image Generation | arXiv | 2023 11-29 | Code Paper |
| Portrait Diffusion: Training-Free Face Stylization with Chain-of-Painting | arXiv | 2023 12-03 | Code Paper |
| Retrieving Conditions from Reference Images for Diffusion Models | arXiv | 2023 12-05 | โ Paper |
| FaceStudio: Put Your Face Everywhere in Seconds | arXiv | 2023 12-05 | Code Paper |
| Personalized Face Inpainting with Diffusion Models by Parallel Visual Attention | WACV 2024 | 2023 12-06 | Code Paper |
| PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding | arXiv | 2023 12-07 | Code Paper |
| DemoCaricature: Democratising Caricature Generation with a Rough Sketch | CVPR 2024 | 2023 12-07 | Code Paper |
| Stellar: Systematic Evaluation of Human-Centric Personalized Text-to-Image Methods | CoRR 2023 | 2023 12-11 | Code Paper |
| PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved Personalization | CVPR 2024 | 2023 12-11 | Code Paper |
| Concept-Centric Personalization with Large-Scale Diffusion Priors | arXiv | 2023 12-13 | Code Paper |
| Cross Initialization for Personalized Text-to-Image Generation | arXiv | 2023 12-26 | Code Paper |
| InstantID: Zero-Shot Identity-Preserving Generation in Seconds | arXiv | 2024 01-15 | Code Paper |
| Face2Diffusion for Fast and Editable Face Personalization | arXiv | 2024 03-08 | โ Paper |
| OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models | arXiv | 2024 03-16 | Code Paper |
| Infinite-ID: Identity-preserved Personalization via ID-semantics Decoupling Paradigm | arXiv | 2024 03-18 | โ Paper |
| IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models | arXiv | 2024 03-21 | โ Paper |
| MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation | arXiv | 2024 04-17 | โ Paper |
| ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning | arXiv | 2024 04-23 | โ Paper |
| InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation | arXiv | 2024 04-30 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Training-free layout control with cross-attention guidance | WACV 2024 | 2023 04-06 | Code Paper |
| Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion Models | CVPR 2024 | 2023 05-25 | Code Paper |
| Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models | NeurIPS 2023 | 2023 05-25 | Code Paper |
| PhotoSwap: Personalized Subject Swapping in Images | NIPS 2023 | 2023 05-29 | Code Paper |
| TryonDiffusion: A Tale of Two UNets | CVPR 2023 | 2023 06-14 | Code Paper |
| ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet | arXiv | 2023 12-05 | Code Paper |
| Context Diffusion: In-Context Aware Image Generation | arXiv | 2023 12-06 | Code Paper |
| FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition | CVPR 2024 | 2023 12-12 | Code Paper |
| A Two-Stage Personalized Virtual Try-On Framework with Shape Control and Texture Guidance | CoRR 2023 | 2023 12-24 | โ Paper |
| Tuning-Free Image Customization with Image and Text Guidance | arXiv | 2024 03-19 | โ Paper |
| SWAPANYTHING: Enabling Arbitrary Object Swapping in Personalized Visual Editing | arXiv | 2024 04-08 | โ Paper |
| Customizing Text-to-Image Diffusion with Camera Viewpoint Control | arXiv | 2024 04-18 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation | ICCV 2023 | 2022 12-22 | Code Paper |
| Structure and Content-Guided Video Synthesis with Diffusion Models | ICCV 2023 | 2023 02-06 | โ Paper |
| Make-A-Protagonist: Generic Video Editing with Visual and Textual Clues | arXiv | 2023 05-15 | Code Paper |
| Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation | arXiv | 2023 07-13 | Code Paper |
| MotionDirector: Motion Customization of Text-to-Video Diffusion Models | arXiv | 2023 10-12 | Code Paper |
| LAMP: Learn a Motion Pattern for Few-Shot-Based Video Generation | arXiv | 2023 10-16 | Code Paper |
| VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning | arXiv | 2023 11-02 | Code Paper |
| VideoAssembler: Identity-Consistent Video Generation with Reference Entities Using Diffusion Model | arXiv | 2023 11-29 | Code Paper |
| VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models | arXiv | 2023 12-01 | Code Paper |
| VideoBooth: Diffusion-Based Video Generation with Image Prompts | CVPR 2024 | 2023 12-01 | Code Paper |
| StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter | arXiv | 2023 12-01 | Code Paper |
| SAVE: Protagonist Diversification with Structure Agnostic Video Editing | arXiv | 2023 12-05 | Code Paper |
| Customizing Motion in Text-to-Video Diffusion Models | arXiv | 2023 12-07 | Code Paper |
| DreamVideo: Composing Your Dream Videos with Customized Subject and Motion | arXiv | 2023 12-07 | Code Paper |
| MotionCrafter: One-Shot Motion Customization of Diffusion Models | arXiv | 2023 12-08 | Code Paper |
| DreaMoving: A Human Video Generation Framework Based on Diffusion Models | arXiv | 2023 12-08 | Code Paper |
| CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects | arXiv | 2024 01-18 | Code Paper |
| Magic-Me: Identity-Specific Video Customized Diffusion | arXiv | 2024 02-14 | โ Paper |
| Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models | arXiv | 2024 02-22 | โ Paper |
| ID-Animator: Zero-Shot Identity-Preserving Human Video Generation | arXiv | 2024 04-23 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Magic3D: High-Resolution Text-to-3D Content Creation | CVPR 2023 | 2022 11-18 | Code Paper |
| DreamBooth3D: Subject-Driven Text-to-3D Generation | ICCV 2023 | 2023 03-23 | Code Paper |
| Text-Conditional Contextualized Avatars For Zero-Shot Personalization | arXiv | 2023 04-14 | โ Paper |
| StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation | arXiv | 2023 05-30 | โ Paper |
| AvatarBooth: High-Quality and Customizable 3D Human Avatar Generation | arXiv | 2023 06-16 | Code Paper |
| MVDREAM: MULTI-VIEW DIFFUSION FOR 3D GENERATION | ICLR 2024 | 2023 08-31 | Code Paper |
| Chasing Consistency in Text-to-3D Generation from a Single Image | arXiv | 2023 09-07 | โ Paper |
| Animate124: Animating One Image to 4D Dynamic Scene | arXiv | 2023 11-24 | Code Paper |
| A Unified Approach for Text- and Image-guided 4D Scene Generation | arXiv | 2023 11-28 | Code Paper |
| TextureDreamer: Image-guided Texture Synthesis through Geometry-aware Diffusion | arXiv | 2024 01-17 | โ Paper |
| TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts | arXiv | 2024 01-26 | Code Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Anti-DreamBooth: Protecting Users from Personalized Text-to-Image Synthesis | arXiv | 2023 03-27 | Code Paper |
| Backdooring Textual Inversion for Concept Censorship | arXiv | 2023 08-21 | Code Paper |
| Personalization as a Shortcut for Few-Shot Backdoor Attack against Text-to-Image Diffusion Models | AAAI | 2024 03-24 | โ Paper |
| ReVersion: Diffusion-Based Relation Inversion from Images | arXiv | 2023 03-23 | Code Paper |
| Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation | arXiv | 2023 11-30 | Code Paper |
| Inv-ReVersion: Enhanced Relation Inversion Based on Text-to-Image Diffusion Models | MDPI | 2024 04-15 | โ Paper |
| Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA | arXiv | 2023 04-12 | Code Paper |
| Text-Guided Vector Graphics Customization | SIGGRAPH 2023 | 2023 09-21 | Code Paper |
| Customizing 360-Degree Panoramas Through Text-to-Image Diffusion Models | WACV 2024 | 2023 10-28 | Code Paper |
If you find any missing work, please report it by creating an Issue in the repository to contribute the community together.
Python
75.1%
Jupyter Notebook
23.6%
Shell
1.3%

This repository contains a collection of papers and resources on Personalized Content Synthesis (PCS) with Diffusion Model.
If you find the information in our paper useful for your research, please consider citing it in your work. Thank you!
@misc{zhang2024survey,
title={A Survey on Personalized Content Synthesis with Diffusion Models},
author={Xulu Zhang and Xiao-Yong Wei and Wengyu Zhang and Jinlin Wu and Zhaoxiang Zhang and Zhen Lei and Qing Li},
year={2024},
eprint={2405.05538},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
To uniformly evaluate Personalized Content Synthesis (PCS) tasks, we introduces a comprehensive evaluation dataset designed for the most common personalized generation tasks, object and face personalization.
Download Link: PCS-dataset
We evaluate existing representative PCS methods based our unified test dataset. The evaluation results and settings are shown in the below table. For more details, please see our survey paper.
| Type | Methods | Framework | Backbone | CLIP-T | CLIP-I |
|---|---|---|---|---|---|
| Object | Textual Inversion | TTF | SD 1.5 | 0.199 | 0.749 |
| Dreambooth | TTF | SD 1.5 | 0.286 | 0.772 | |
| P+ | TTF | SD 1.4 | 0.244 | 0.643 | |
| Custom Diffusion | TTF | SD 1.4 | 0.307 | 0.722 | |
| NeTI | TTF | SD 1.4 | 0.283 | 0.801 | |
| SVDiff | TTF | SD 1.5 | 0.282 | 0.776 | |
| Perfusion | TTF | SD 1.5 | 0.273 | 0.691 | |
| ELITE | PTA | SD 1.4 | 0.292 | 0.765 | |
| BLIP-Diffusion | PTA | SD 1.5 | 0.292 | 0.772 | |
| IP-Adapter | PTA | SD 1.5 | 0.272 | 0.825 | |
| SSR Encoder | PTA | SD 1.5 | 0.288 | 0.792 | |
| MoMA | PTA | SD 1.5 | 0.322 | 0.748 | |
| Diptych Prompting | PTA | FLUX 1.0 dev | 0.327 | 0.722 | |
| ฮป-eclipse | PTA | Kandinsky 2.2 | 0.272 | 0.824 | |
| MS-Diffusion | PTA | SDXL | 0.298 | 0.777 | |
| Face | CrossInitialization | TTF | SD 2.1 | 0.261 | 0.469 |
| Face2Diffusion | PTA | SD 1.4 | 0.265 | 0.588 | |
| SSR Encoder | PTA | SD 1.5 | 0.233 | 0.490 | |
| FastComposer | PTA | SD 1.5 | 0.230 | 0.516 | |
| IP-Adapter | PTA | SD 1.5 | 0.292 | 0.462 | |
| IP-Adapter | PTA | SDXL | 0.292 | 0.642 | |
| PhotoMaker | PTA | SDXL | 0.311 | 0.547 | |
| InstantID | PTA | SDXL | 0.278 | 0.707 |
TTF: Test-time Fine-tuning, PTA: Pre-trained Adaptation
| Title | Venue | Date | Links |
|---|---|---|---|
| An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion | ICLR 2023 | 2022 08-02 | Code Paper |
| DreamBooth: Fine-Tuning Text-to-Image Diffusion Models for Subject-Driven Generation | CVPR 2023 | 2022 08-25 | Code Paper |
| Re-Imagen: Retrieval-Augmented Text-to-Image Generator | ICLR 2023 | 2022 09-29 | โ Paper |
| Versatile Diffusion: Text, Images, and Variations All in One Diffusion Model | ICCV 2023 | 2022 11-15 | Code Paper |
| DreamArtist: Towards Controllable One-Shot Text-to-Image Generation via Positive-Negative Prompt-Tuning | arXiv 2022 | 2022 11-21 | Code Paper |
| Is This Loss Informative? Faster Text-to-Image Customization by Tracking Objective Dynamics | NeurIPS 2023 | 2023 02-09 | Code Paper |
| Encoder-Based Domain Tuning for Fast Personalization of Text-to-Image Models | ACM Trans on Graphics | 2023 02-23 | Code Paper |
| ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation | ICCV 2023 | 2023 02-27 | Code Paper |
| Highly Personalized Text Embedding for Image Manipulation by Stable Diffusion | arXiv | 2023 03-15 | Code Paper |
| Unified Multi-Modal Latent Diffusion for Joint Subject and Text Conditional Image Generation | arXiv | 2023 03-16 | โ Paper |
| P+: Extended Textual Conditioning in Text-to-Image Generation | arXiv | 2023 03-16 | Code Paper |
| A Closer Look at Parameter-Efficient Tuning in Diffusion Models | arXiv | 2023 03-31 | Code Paper |
| Subject-Driven Text-to-Image Generation via Apprenticeship Learning | NIPS 2023 | 2023 04-01 | โ Paper |
| Taming Encoder for Zero Fine-Tuning Image Customization with Text-to-Image Diffusion Models | arXiv | 2023 04-05 | โ Paper |
| InstantBooth: Personalized Text-to-Image Generation Without Test-Time Finetuning | arXiv | 2023 04-06 | Code Paper |
| Controllable Textual Inversion for Personalized Text-to-Image Generation | arXiv | 2023 04-11 | Code Paper |
| Gradient-Free Textual Inversion | ACM MM 2023 | 2023 04-12 | Code Paper |
| Personalize Segment Anything Model with One Shot | ICLR 2024 | 2023 05-04 | Code Paper |
| DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation | ICLR 2024 | 2023 05-05 | Code Paper |
| BLIP-Diffusion: Pre-Trained Subject Representation for Controllable Text-to-Image Generation and Editing | NIPS 2023 | 2023 05-24 | Code Paper |
| A Neural Space-Time Representation for Text-to-Image Personalization | SIGGRAPH Asia 2023 | 2023 05-24 | Code Paper |
| Prospect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models | ACM Trans on Graphics | 2023 05-25 | Code Paper |
| Break-a-Scene: Extracting Multiple Concepts from a Single Image | SIGGRAPH ASIA 2023 | 2023 05-25 | Code Paper |
| COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models | arXiv | 2023 05-26 | Code Paper |
| ViCo: Plug-and-Play Visual Condition for Personalized Text-to-Image Generation | arXiv | 2023 06-01 | Code Paper |
| Controlling Text-to-Image Diffusion by Orthogonal Fine-Tuning | arXiv | 2023 06-12 | Code Paper |
| Domain-Agnostic Tuning-Encoder for Fast Personalization of Text-to-Image Models | SIGGRAPH 2023 | 2023 07-13 | Code Paper |
| IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models | arXiv | 2023 08-13 | Code Paper |
| Navigating Text-to-Image Customization: From Lycoris Fine-Tuning to Model Evaluation | ICLR 2024 | 2023 09-26 | Code Paper |
| Kosmos-G: Generating Images in Context with Multimodal Large Language Models | arXiv | 2023 10-04 | Code Paper |
| Personalized Text-to-Image Model Enhancement Strategies: SOD Preprocessing and CNN Local Feature Integration | arXiv | 2023 10-26 | โ Paper |
| A Data Perspective on Enhanced Identity Preservation for Diffusion Personalization | ICLR 2024 | 2023 11-07 | Code Paper |
| DIFFNAT: Improving Diffusion Image Quality Using Natural Image Statistics | arXiv | 2023 11-16 | โ Paper |
| An Image is Worth Multiple Words: Multi-attribute Inversion for Constrained Text-to-Image Synthesis | arXiv | 2023 11-20 | โ Paper |
| LEGO: Learning to Disentangle and Invert Concepts Beyond Object Appearance in Text-to-Image Diffusion Models | arXiv | 2023 11-23 | Code Paper |
| Catversion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization | arXiv | 2023 11-24 | Code Paper |
| CLiC: Concept Learning in Context | CVPR 2024 | 2023 11-28 | Code Paper |
| HiFi Tuner: High-Fidelity Subject-Driven Fine-Tuning for Diffusion Models | arXiv | 2023 11-30 | โ Paper |
| InstructBooth: Instruction-Following Personalized Text-to-Image Generation | arXiv | 2023 12-04 | โ Paper |
| Customization Assistant for Text-to-Image Generation | CVPR 2024 | 2023 12-05 | โ Paper |
| Decoupled Textual Embeddings for Customized Image Generation | AAAI 2024 | 2023 12-19 | Code Paper |
| Towards Accurate Guided Diffusion Sampling through Symplectic Adjoint Method | arXiv | 2023 12-19 | Code Paper |
| DreamDistribution: Prompt Distribution Learning for Text-to-Image Diffusion Models | arXiv | 2023 12-21 | Code Paper |
| DreamTuner: Single Image is Enough for Subject-Driven Generation | arXiv | 2023 12-21 | Code Paper |
| BootPIG: Bootstrapping Zero-Shot Personalized Image Generation Capabilities in Pretrained Diffusion Models | arXiv | 2024 01-25 | โ Paper |
| Object-Driven One-Shot Fine-Tuning of Text-to-Image Diffusion with Prototypical Embedding | arXiv | 2024 01-28 | โ Paper |
| DisenDreamer: Subject-Driven Text-to-Image Generation with Sample-aware Disentangled Tuning | arXiv | 2024 02-26 | โ Paper |
| Infusion: Preventing Customized Text-to-Image Diffusion from Overfitting | arXiv | 2024 04-22 | โ Paper |
| Customizing Text-to-Image Models with a Single Image Pair | arXiv | 2024 05-02 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Multi-Concept Customization of Text-to-Image Diffusion | CVPR 2023 | 2022 12-08 | Code Paper |
| CONES: Concept Neurons in Diffusion Models for Customized Generation | ICML 2023 | 2023 03-09 | Code Paper |
| SVDiff: Compact Parameter Space for Diffusion Fine-Tuning | ICCV 2023 | 2023 03-20 | Code Paper |
| Key-Locked Rank One Editing for Text-to-Image Personalization | SIGGRAPH 2023 | 2023 05-02 | โ Paper |
| Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models | NIPS 2023 | 2023 05-29 | Code Paper |
| CONES 2: Customizable Image Synthesis with Multiple Subjects | NIPS 2023 | 2023 05-30 | Code Paper |
| Generate Anything Anywhere in Any Scene | arXiv | 2023 06-29 | Code Paper |
| AnyDoor: Zero-Shot Object-Level Image Customization | arXiv | 2023 07-18 | Code Paper |
| Subject-Diffusion: Open Domain Personalized Text-to-Image Generation Without Test-Time Fine-Tuning | arXiv | 2023 07-21 | Code Paper |
| CustomNet: Zero-Shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models | arXiv | 2023 10-30 | Code Paper |
| Compositional Inversion for Stable Diffusion Models | AAAI 2024 | 2023 12-13 | Code Paper |
| Visual Concept-Driven Image Generation with Text-to-Image Diffusion Model | arXiv | 2024 02-18 | โ Paper |
| MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis | arXiv | 2024 02-27 | โ Paper |
| Multi-Object Editing in Personalized Text-To-Image Diffusion Model Via Segmentation Guidance | arXiv | 2024 03-18 | โ Paper |
| MC2: Multi-concept Guidance for Customized Multi-concept Generation | arXiv | 2024 04-12 | โ Paper |
| MultiBooth: Towards Generating All Your Concepts in an Image from Text | arXiv | 2024 04-22 | โ Paper |
| MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models | arXiv | 2024 10-06 | Code Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| StyleDrop: Text-to-Image Synthesis of Any Style | NIPS 2023 | 2023 06-01 | Code Paper |
| StyleAdapter: A Single-Pass LoRA-Free Model for Stylized Image Generation | ICLR 2024 | 2023 09-04 | โ Paper |
| StyleBoost: A Study of Personalizing Text-to-Image Generation in Any Style using DreamBooth | ICTC 2023 | 2023 10-13 | โ Paper |
| ArtAdapter: Text-to-Image Style Transfer Using Multi-Level Style Encoder and Explicit Adaptation | arXiv | 2023 12-04 | Code Paper |
| Style Aligned Image Generation via Shared Attention | CVPR 2024 | 2023 12-04 | Code Paper |
| Generative Active Learning for Image Synthesis Personalization | arXiv | 2024 03-22 | Code Paper |
| Text-to-Image Synthesis for Any Artistic Styles: Advancements in Personalized Artistic Image Generation via Subdivision and Dual Binding | arXiv | 2024 04-08 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Identity Encoder for Personalized Diffusion | CoRR 2023 | 2023 04-14 | โ Paper |
| FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention | CoRR 2023 | 2023 05-21 | Code Paper |
| Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach | arXiv | 2023 05-23 | โ Paper |
| Inserting Anybody in Diffusion Models via Celeb Basis | NIPS 2023 | 2023 06-01 | Code Paper |
| Face0: Instantaneously Conditioning a Text-to-Image Model on a Face | SIGGRAPH 2023 | 2023 06-11 | โ Paper |
| DreamIdentity: Improved Editability for Efficient Face-Identity Preserved Image Generation | arXiv | 2023 07-01 | Code Paper |
| HyperDreamBooth: Hypernetworks for Fast Personalization of Text-to-Image Models | arXiv | 2023 07-13 | Code Paper |
| Identity-Preserving Aging of Face Images via Latent Diffusion Models | IJCB 2023 | 2023 07-17 | Code Paper |
| Magicapture: High-Resolution Multi-Concept Portrait Customization | arXiv | 2023 09-13 | Code Paper |
| High-Fidelity Person-Centric Subject-to-Image Synthesis | CVPR 2024 | 2023 11-17 | โ Paper |
| When StyleGAN Meets Stable Diffusion: A W+ Adapter for Personalized Image Generation | arXiv | 2023 11-29 | Code Paper |
| Portrait Diffusion: Training-Free Face Stylization with Chain-of-Painting | arXiv | 2023 12-03 | Code Paper |
| Retrieving Conditions from Reference Images for Diffusion Models | arXiv | 2023 12-05 | โ Paper |
| FaceStudio: Put Your Face Everywhere in Seconds | arXiv | 2023 12-05 | Code Paper |
| Personalized Face Inpainting with Diffusion Models by Parallel Visual Attention | WACV 2024 | 2023 12-06 | Code Paper |
| PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding | arXiv | 2023 12-07 | Code Paper |
| DemoCaricature: Democratising Caricature Generation with a Rough Sketch | CVPR 2024 | 2023 12-07 | Code Paper |
| Stellar: Systematic Evaluation of Human-Centric Personalized Text-to-Image Methods | CoRR 2023 | 2023 12-11 | Code Paper |
| PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved Personalization | CVPR 2024 | 2023 12-11 | Code Paper |
| Concept-Centric Personalization with Large-Scale Diffusion Priors | arXiv | 2023 12-13 | Code Paper |
| Cross Initialization for Personalized Text-to-Image Generation | arXiv | 2023 12-26 | Code Paper |
| InstantID: Zero-Shot Identity-Preserving Generation in Seconds | arXiv | 2024 01-15 | Code Paper |
| Face2Diffusion for Fast and Editable Face Personalization | arXiv | 2024 03-08 | โ Paper |
| OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models | arXiv | 2024 03-16 | Code Paper |
| Infinite-ID: Identity-preserved Personalization via ID-semantics Decoupling Paradigm | arXiv | 2024 03-18 | โ Paper |
| IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models | arXiv | 2024 03-21 | โ Paper |
| MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation | arXiv | 2024 04-17 | โ Paper |
| ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning | arXiv | 2024 04-23 | โ Paper |
| InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation | arXiv | 2024 04-30 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Training-free layout control with cross-attention guidance | WACV 2024 | 2023 04-06 | Code Paper |
| Prompt-Free Diffusion: Taking "Text" Out of Text-to-Image Diffusion Models | CVPR 2024 | 2023 05-25 | Code Paper |
| Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models | NeurIPS 2023 | 2023 05-25 | Code Paper |
| PhotoSwap: Personalized Subject Swapping in Images | NIPS 2023 | 2023 05-29 | Code Paper |
| TryonDiffusion: A Tale of Two UNets | CVPR 2023 | 2023 06-14 | Code Paper |
| ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet | arXiv | 2023 12-05 | Code Paper |
| Context Diffusion: In-Context Aware Image Generation | arXiv | 2023 12-06 | Code Paper |
| FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition | CVPR 2024 | 2023 12-12 | Code Paper |
| A Two-Stage Personalized Virtual Try-On Framework with Shape Control and Texture Guidance | CoRR 2023 | 2023 12-24 | โ Paper |
| Tuning-Free Image Customization with Image and Text Guidance | arXiv | 2024 03-19 | โ Paper |
| SWAPANYTHING: Enabling Arbitrary Object Swapping in Personalized Visual Editing | arXiv | 2024 04-08 | โ Paper |
| Customizing Text-to-Image Diffusion with Camera Viewpoint Control | arXiv | 2024 04-18 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation | ICCV 2023 | 2022 12-22 | Code Paper |
| Structure and Content-Guided Video Synthesis with Diffusion Models | ICCV 2023 | 2023 02-06 | โ Paper |
| Make-A-Protagonist: Generic Video Editing with Visual and Textual Clues | arXiv | 2023 05-15 | Code Paper |
| Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation | arXiv | 2023 07-13 | Code Paper |
| MotionDirector: Motion Customization of Text-to-Video Diffusion Models | arXiv | 2023 10-12 | Code Paper |
| LAMP: Learn a Motion Pattern for Few-Shot-Based Video Generation | arXiv | 2023 10-16 | Code Paper |
| VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning | arXiv | 2023 11-02 | Code Paper |
| VideoAssembler: Identity-Consistent Video Generation with Reference Entities Using Diffusion Model | arXiv | 2023 11-29 | Code Paper |
| VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models | arXiv | 2023 12-01 | Code Paper |
| VideoBooth: Diffusion-Based Video Generation with Image Prompts | CVPR 2024 | 2023 12-01 | Code Paper |
| StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter | arXiv | 2023 12-01 | Code Paper |
| SAVE: Protagonist Diversification with Structure Agnostic Video Editing | arXiv | 2023 12-05 | Code Paper |
| Customizing Motion in Text-to-Video Diffusion Models | arXiv | 2023 12-07 | Code Paper |
| DreamVideo: Composing Your Dream Videos with Customized Subject and Motion | arXiv | 2023 12-07 | Code Paper |
| MotionCrafter: One-Shot Motion Customization of Diffusion Models | arXiv | 2023 12-08 | Code Paper |
| DreaMoving: A Human Video Generation Framework Based on Diffusion Models | arXiv | 2023 12-08 | Code Paper |
| CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects | arXiv | 2024 01-18 | Code Paper |
| Magic-Me: Identity-Specific Video Customized Diffusion | arXiv | 2024 02-14 | โ Paper |
| Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models | arXiv | 2024 02-22 | โ Paper |
| ID-Animator: Zero-Shot Identity-Preserving Human Video Generation | arXiv | 2024 04-23 | โ Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Magic3D: High-Resolution Text-to-3D Content Creation | CVPR 2023 | 2022 11-18 | Code Paper |
| DreamBooth3D: Subject-Driven Text-to-3D Generation | ICCV 2023 | 2023 03-23 | Code Paper |
| Text-Conditional Contextualized Avatars For Zero-Shot Personalization | arXiv | 2023 04-14 | โ Paper |
| StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation | arXiv | 2023 05-30 | โ Paper |
| AvatarBooth: High-Quality and Customizable 3D Human Avatar Generation | arXiv | 2023 06-16 | Code Paper |
| MVDREAM: MULTI-VIEW DIFFUSION FOR 3D GENERATION | ICLR 2024 | 2023 08-31 | Code Paper |
| Chasing Consistency in Text-to-3D Generation from a Single Image | arXiv | 2023 09-07 | โ Paper |
| Animate124: Animating One Image to 4D Dynamic Scene | arXiv | 2023 11-24 | Code Paper |
| A Unified Approach for Text- and Image-guided 4D Scene Generation | arXiv | 2023 11-28 | Code Paper |
| TextureDreamer: Image-guided Texture Synthesis through Geometry-aware Diffusion | arXiv | 2024 01-17 | โ Paper |
| TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts | arXiv | 2024 01-26 | Code Paper |
| Title | Venue | Date | Links |
|---|---|---|---|
| Anti-DreamBooth: Protecting Users from Personalized Text-to-Image Synthesis | arXiv | 2023 03-27 | Code Paper |
| Backdooring Textual Inversion for Concept Censorship | arXiv | 2023 08-21 | Code Paper |
| Personalization as a Shortcut for Few-Shot Backdoor Attack against Text-to-Image Diffusion Models | AAAI | 2024 03-24 | โ Paper |
| ReVersion: Diffusion-Based Relation Inversion from Images | arXiv | 2023 03-23 | Code Paper |
| Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation | arXiv | 2023 11-30 | Code Paper |
| Inv-ReVersion: Enhanced Relation Inversion Based on Text-to-Image Diffusion Models | MDPI | 2024 04-15 | โ Paper |
| Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA | arXiv | 2023 04-12 | Code Paper |
| Text-Guided Vector Graphics Customization | SIGGRAPH 2023 | 2023 09-21 | Code Paper |
| Customizing 360-Degree Panoramas Through Text-to-Image Diffusion Models | WACV 2024 | 2023 10-28 | Code Paper |
If you find any missing work, please report it by creating an Issue in the repository to contribute the community together.
Python
75.1%
Jupyter Notebook
23.6%
Shell
1.3%