neptune-T/Awesome-Style-Transfer

A comprehensive collection of papers and datasets on generative models and their applications in style transfer across image, text, 3D, and video domains.

47

81 commits

updated Aug 11, 2026

See the code

README

🎨 Awesome Style Transfer

Awesome Papers Carriers Update License: MIT Stars

Taxonomy | Papers | Coverage | Datasets | Contributing

🎨 About This Repository

A curated collection of style transfer research papers spanning 🖼️ 2D image, 🎬 video, 🧊 3D (NeRF / mesh / point cloud / 3DGS) and ⏳ 4D (dynamic scene) stylization — from classic neural style transfer to the latest diffusion- and foundation-model-era methods — organized by the Style-Carrier Taxonomy from our survey Style Transfer: A Decade Survey (revised version coming soon).

Domain tags: 🖼️ Image — single-image stylization · 🎬 Video — temporally consistent stylization · 🧊 3D — stylization of neural / explicit 3D representations · ⏳ 4D — stylization of dynamic, time-varying 3D scenes · 🌐 Multiple — methods covering more than one domain

Instead of grouping papers by year, backbone (CNN / GAN / Diffusion / DiT), or output domain, we ask a single question:

In what form is style-specific information materialized, and how is it reused for new content?

Every method has exactly one canonical home — five style carriers spanning 2015–2026 across 🖼️ image, 🎬 video, 🧊 3D and ⏳ 4D stylization.

🖼️ Style Transfer Examples

Content ImageStyle ImageTransferred Result
Disaster GirlとなりのトトロTransferred Result

🧭 The Style-Carrier Taxonomy

CarrierCore ideaAnchorsPapers
🎯 Style as Optimization ObjectiveStyle exists only as a loss, score, or guidance signal — no transferable style representation is ever materialized.Gatys et al. 201520
🧩 Style as Feature OperatorStyle acts through direct matching, replacement, transport, or attention coupling of content/style activations.AdaIN · WCT27
🌀 Style as Latent VariableStyle is a native generative latent with its own prior — sampleable, mixable, and editable without any reference.MUNIT · StyleGAN11
⚙️ Style as Model ParametersStyle is materialized as a persistent parameter asset (weights, LoRA, learned tokens) reusable across new content.Johnson 2016 · LoRA25
🔌 Style as Learned ConditionAn external reference is encoded into a standalone style condition consumed via a dedicated conditioning interface.Ghiasi 2017 · IP-Adapter27

📌 What the taxonomy is not based on: year, backbone family, output domain, or auxiliary mechanisms (ControlNet, SDS, DDIM inversion …) — those are tags, not categories. The same perceptual objective underlies Gatys (🎯 objective), Johnson (⚙️ parameter) and Ghiasi (🔌 condition); AdaIN, StyTr² and StyleAligned are all 🧩 feature operators despite spanning CNN, Transformer and Diffusion eras.

📊 Coverage at a Glance

Carrier🖼️ Image🎬 Video🧊 3D⏳ 4D
🎯 Optimization Objective6—121
🧩 Feature Operator2221—
🌀 Latent Variable92——
⚙️ Model Parameters1861—
🔌 Learned Condition2123—

Empty cells are open problems, not omissions — e.g. ⏳ 4D stylization remains nearly untouched.

📑 Table of Contents

  1. 🎯 Style as Optimization Objective — 20 papers
  2. 🧩 Style as Feature Operator — 27 papers
  3. 🌀 Style as Latent Variable — 11 papers
  4. ⚙️ Style as Model Parameters — 25 papers
  5. 🔌 Style as Learned Condition — 27 papers
  6. 📚 Datasets
  7. 🤝 Contributing

🗂️ Papers by Style Carrier

1. 🎯 Style as Optimization Objective

Style exists only as a loss, score, or guidance signal — no transferable style representation is ever materialized.

PaperVenueYearRepLink
A neural algorithm of artistic stylearXiv2015🖼️paper
Deep photo style transferCVPR2017🖼️paper
Style transfer by relaxed optimal transport and self-similarityCVPR2019🖼️paper
Arf: Artistic radiance fieldsECCV2022🧊paper
CLIPstyler: Image style transfer with a single text conditionCVPR2022🖼️paper
SNeRF: Stylized neural implicit representations for 3D scenesSIGGRAPH Asia2022🧊paper
StyleMesh: Style transfer for indoor 3D scene reconstructionsCVPR2022🧊paper
Diffusion-based image translation using disentangled style and content representationICLR2023🖼️paper
Ref-NPR: Reference-Based Non-Photorealistic Radiance Fields for Controllable Scene StylizationCVPR2023🧊paper
3D Paintbrush: Local Stylization of 3D Shapes with Cascaded Score DistillationCVPR2024🧊paper
NeRF-Art: Text-Driven Neural Radiance Fields StylizationIEEE TVCG2024🧊paper
Balanced image stylization with style matching scoreICCV2025🖼️paper
CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian SplattingNeurIPS2025🌐paper
GAS-NeRF: Geometry-Aware Stylization of Dynamic Radiance FieldsarXiv2025⏳paper
GT²-GS: Geometry-aware Texture Transfer for Gaussian SplattingarXiv2025🧊paper
Morpheus: Text-Driven 3D Gaussian Splat Shape and Color StylizationCVPR2025🧊paper
SGSST: Scaling Gaussian Splatting Style TransferCVPR2025🧊paper
Stylizedgs: Controllable stylization for 3d gaussian splattingTPAMI2025🧊paper
DiffStyle3D: Consistent 3D Gaussian Stylization via Attention OptimizationarXiv2026🧊paper
Fantasystyle: Controllable stylized distillation for 3d gaussian splattingAAAI2026🧊paper

2. 🧩 Style as Feature Operator

Style acts through direct matching, replacement, transport, or attention coupling of content/style activations.

PaperVenueYearRepLink
Arbitrary style transfer in real-time with adaptive instance normalizationICCV2017🖼️paper
Universal style transfer via feature transformsNeurIPS2017🖼️paper
A closed-form solution to photorealistic image stylizationECCV2018🖼️paper
Avatar-net: Multi-scale zero-shot style transfer by feature decorationCVPR2018🖼️paper
Arbitrary Style Transfer with Style-Attentional NetworksCVPR2019🖼️paper
Learning linear transformations for fast image and video style transferCVPR2019🌐paper
Photorealistic style transfer via wavelet transformsICCV2019🖼️paper
AdaAttN: Revisit attention mechanism in arbitrary neural style transferICCV2021🖼️paper
Arbitrary video style transfer via multi-channel correlationICCV2021🎬paper
Artflow: Unbiased image style transfer via reversible neural flowsCVPR2021🖼️paper
CCPL: Contrastive coherence preserving loss for versatile style transferECCV2022🌐paper
Domain enhanced arbitrary image style transfer via contrastive learningSIGGRAPH2022🖼️paper
Exact feature distribution matching for arbitrary style transfer and domain generalizationCVPR2022🖼️paper
Stytr2: Image style transfer with transformersCVPR2022🖼️paper
CAP-VSTNet: Content affinity preserved versatile style transferCVPR2023🖼️paper
Ctrl-x: Controlling structure and appearance for text-to-image generation without guidanceNeurIPS2024🖼️paper
Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion NetworkCVPR2024🖼️paper
Style aligned image generation via shared attentionCVPR2024🖼️paper
Style injection in diffusion: A training-free approach for adapting large-scale diffusion models for style transferCVPR2024🖼️paper
Training-free consistent text-to-image generationACM TOG2024🖼️paper
AttenST: A Training-Free Attention-Driven Style Transfer Framework with Pre-Trained Diffusion ModelsarXiv2025🖼️paper
DVI: Disentangling Semantic and Visual Identity for Training-Free Personalized GenerationarXiv2025🖼️paper
FPGS: Feed-Forward Semantic-aware Photorealistic Style Transfer of Large-Scale Gaussian SplattingInt. J. Comput. Vis. 20262025🧊paper
SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style TransferCVPR2025🖼️paper
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion ModelsarXiv2025🎬paper
StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style TransferCVPR2025🖼️paper
StyleStudio: Text-Driven Style Transfer with Selective Control of Style ElementsCVPR2025🖼️paper

3. 🌀 Style as Latent Variable

Style is a native generative latent with its own prior — sampleable, mixable, and editable without any reference.

PaperVenueYearRepLink
Diverse image-to-image translation via disentangled representationsECCV2018🖼️paper
MoCoGAN: Decomposing motion and content for video generationCVPR2018🎬paper
Multimodal unsupervised image-to-image translationECCV2018🖼️paper
Multimodal Unsupervised Image-to-image TranslationECCV2018🖼️paper
A style-based generator architecture for generative adversarial networksCVPR2019🖼️paper
Analyzing and Improving the Image Quality of StyleGANCVPR2020🖼️paper
Analyzing and Improving the Image Quality of StyleGANCVPR2020🖼️paper
Alias-Free Generative Adversarial NetworksNeurIPS2021🖼️paper
Styleclip: Text-driven manipulation of stylegan imageryICCV2021🖼️paper
Diffusion autoencoders: Toward a meaningful and decodable representationCVPR2022🖼️paper
Styleinv: A temporal style modulated inversion network for unconditional video generationICCV2023🎬paper

4. ⚙️ Style as Model Parameters

Style is materialized as a persistent parameter asset (weights, LoRA, learned tokens) reusable across new content.

PaperVenueYearRepLink
Perceptual losses for real-time style transfer and super-resolutionECCV2016🖼️paper
Texture networks: Feed-forward synthesis of textures and stylized imagesICML2016🖼️paper
A Learned Representation for Artistic StyleICLR2017🖼️paper
Coherent online video style transferICCV2017🎬paper
Real-time neural style transfer for videosCVPR2017🎬paper
ReCoNet: Real-time coherent video style transfer networkECCV2018🎬paper
VToonify: Controllable High-Resolution Portrait Video Style TransferACM TOG2022🎬paper
An image is worth one word: Personalizing text-to-image generation using textual inversionICLR2023🖼️paper
Break-a-scene: Extracting multiple concepts from a single imageSIGGRAPH Asia2023🖼️paper
Cones: Concept neurons in diffusion models for customized generationICML2023🖼️paper
DreamBooth3D: Subject-Driven Text-to-3D GenerationICCV2023🧊paper
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generationCVPR2023🖼️paper
Inversion-Based Style Transfer With Diffusion ModelsCVPR2023🖼️paper
Key-locked rank one editing for text-to-image personalizationSIGGRAPH2023🖼️paper
Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion modelsNeurIPS2023🖼️paper
Multi-concept customization of text-to-image diffusionCVPR2023🖼️paper
Prospect: Prompt spectrum for attribute-aware personalization of diffusion modelsACM TOG2023🖼️paper
Styledrop: Text-to-image generation in any stylearXiv2023🖼️paper
Dreamvideo: Composing your dream videos with customized subject and motionCVPR2024🎬paper
Implicit style-content separation using b-loraECCV2024🖼️paper
Magic-me: Identity-specific video customized diffusionECCV2024🎬paper
ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAsECCV2024🖼️paper
APT: Adaptive Personalized Training for Diffusion Models with Limited DataCVPR2025🖼️paper
ConsisLoRA: Enhancing Content and Style Consistency for LoRA-based Style TransferarXiv2025🖼️paper
Photodoodle: Learning artistic image editing from few-shot pairwise dataarXiv2025🖼️paper

5. 🔌 Style as Learned Condition

An external reference is encoded into a standalone style condition consumed via a dedicated conditioning interface.

PaperVenueYearRepLink
CLIP-NeRF: Text-and-image driven manipulation of neural radiance fieldsCVPR2022🧊paper
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editingNeurIPS2023🖼️paper
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generationICCV2023🖼️paper
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion ModelsarXiv2023🖼️paper
Arc2Face: A Foundation Model for ID-Consistent Human FacesECCV2024🖼️paper
ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationCVPR2024🖼️paper
Deadiff: An efficient stylization diffusion model with disentangled representationsCVPR2024🖼️paper
DiffStyler: Diffusion-based Localized Image Style TransferarXiv2024🖼️paper
Instantid: Zero-shot identity-preserving generation in secondsarXiv2024🖼️paper
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image GenerationarXiv2024🖼️paper
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image GenerationarXiv2024🖼️paper
Photomaker: Customizing realistic human photos via stacked id embeddingCVPR2024🖼️paper
Pulid: Pure and lightning id customization via contrastive alignmentNeurIPS2024🖼️paper
Ssr-encoder: Encoding selective subject representation for subject-driven generationCVPR2024🖼️paper
StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter LearningACM TOG2024🎬paper
CSGO: Content-Style Composition in Text-to-Image GenerationNeurIPS2025🖼️paper
Free-lunch color-texture disentanglement for stylized image generationNeurIPS2025🖼️paper
ID-Consistent, Precise Expression Generation with Blendshape-Guided DiffusionICCV2025🖼️paper
Nested attention: Semantic-aware attention values for concept personalizationSIGGRAPH2025🖼️paper
Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationarXiv2025🖼️paper
Stylemaster: Stylize your video with artistic generation and translationCVPR2025🎬paper
Styleshot: A snapshot on any styleTPAMI2025🖼️paper
Stylos: Multi-View 3D Stylization with Single-Forward Gaussian SplattingarXiv2025🧊paper
U-StyDiT: Ultra-high quality artistic style transfer using diffusion transformersarXiv preprint arXiv:2503.081572025🖼️paper
Uso: Unified style and subject-driven generation via disentangled and reward learningarXiv preprint arXiv:2508.189662025🖼️paper
AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian SplattingarXiv2026🧊paper
TeleStyle: Content-Preserving Style Transfer in Images and VideosarXiv preprint arXiv:2601.201752026🌐paper

📚 Datasets

🖼️ Image Datasets
DatasetYearSizeDescriptionLink
WikiArt201842,129 imagesStyle, Artist, Genrelink
Stylized ImageNet2018~134GBStyle Transferlink
BAM20172.5M imagesArtistic Medialink
ArtBench-10202260,000 imagesArtwork Benchmarklink
FFHQ201970,000 imagesHuman Faceslink
MetFaces20201,336 imagesArtistic Faceslink
AAHQ202125,000 imagesArtistic Faceslink
Ukiyo-e Faces20205,209 imagesAligned Ukiyo-e Faceslink
DiffusionDB202214M imagesText-to-imagelink
JourneyDB20234.4M imagesMultimodal Visionlink
StyleShot2024-Style Transferlink
Danbooru201720172.94M imagesAnimelink
Chinese Style Transfer20181000 content, 100 styleChinese Paintinglink
4SKST202325 color, 100 sketchesSketch Stylelink
🎬 Video Datasets
DatasetYearSizeDescriptionLink
UADFV2018100 videosVideo Style Transferlink
FaceForensics++20196000 videosSwapped Facelink
Celeb-DF2020408 original videosDeepFakelink
DFDC2020100,000 clipsDeepFake Detectionlink
FFIW-10K202110,000 videosFace Forgerylink
ForgeryNet2021221,247 videosForgery Analysislink
🧊 3D / Motion Datasets
DatasetYearSizeDescriptionLink
100STYLE20224M framesStylized Motion Capturelink
Motiondataset202336,673 frames3D Motionlink

🤝 Contributing

We welcome contributions! When suggesting a paper, please indicate which style carrier it belongs to:

  • 🎯 Objective — style lives only in the loss / score / guidance
  • 🧩 Feature — direct activation matching, transport, or attention coupling
  • 🌀 Latent — native generative latent, sampleable without a reference
  • ⚙️ Parameter — persistent style asset (weights / LoRA / tokens)
  • 🔌 Condition — reference-encoded standalone style condition

Ways to help:

  • 🐛 Report bugs and issues
  • 💡 Suggest new papers or resources
  • 🔧 Submit pull requests
  • ⭐ Star this repository if you find it helpful!

📖 Citation

If you find this repository useful for your research, please consider citing:

@article{zhang2025style,
  title={Style transfer: A decade survey},
  author={Zhang, Tianshan and Tang, Hao},
  journal={arXiv preprint arXiv:2506.19278},
  year={2025}
}

⭐ Star History

Star History Chart

🙏 Acknowledgments

Thanks to all researchers and developers who made their work publicly available.

By Monet's Impression of Sunrise

Contributors

neptune-T

79 commits

00lostin00

2 commits

neptune-T/Awesome-Style-Transfer

A comprehensive collection of papers and datasets on generative models and their applications in style transfer across image, text, 3D, and video domains.

47

81 commits

updated Aug 11, 2026

See the code

README

🎨 Awesome Style Transfer

Awesome Papers Carriers Update License: MIT Stars

Taxonomy | Papers | Coverage | Datasets | Contributing

🎨 About This Repository

A curated collection of style transfer research papers spanning 🖼️ 2D image, 🎬 video, 🧊 3D (NeRF / mesh / point cloud / 3DGS) and ⏳ 4D (dynamic scene) stylization — from classic neural style transfer to the latest diffusion- and foundation-model-era methods — organized by the Style-Carrier Taxonomy from our survey Style Transfer: A Decade Survey (revised version coming soon).

Domain tags: 🖼️ Image — single-image stylization · 🎬 Video — temporally consistent stylization · 🧊 3D — stylization of neural / explicit 3D representations · ⏳ 4D — stylization of dynamic, time-varying 3D scenes · 🌐 Multiple — methods covering more than one domain

Instead of grouping papers by year, backbone (CNN / GAN / Diffusion / DiT), or output domain, we ask a single question:

In what form is style-specific information materialized, and how is it reused for new content?

Every method has exactly one canonical home — five style carriers spanning 2015–2026 across 🖼️ image, 🎬 video, 🧊 3D and ⏳ 4D stylization.

🖼️ Style Transfer Examples

Content ImageStyle ImageTransferred Result
Disaster GirlとなりのトトロTransferred Result

🧭 The Style-Carrier Taxonomy

CarrierCore ideaAnchorsPapers
🎯 Style as Optimization ObjectiveStyle exists only as a loss, score, or guidance signal — no transferable style representation is ever materialized.Gatys et al. 201520
🧩 Style as Feature OperatorStyle acts through direct matching, replacement, transport, or attention coupling of content/style activations.AdaIN · WCT27
🌀 Style as Latent VariableStyle is a native generative latent with its own prior — sampleable, mixable, and editable without any reference.MUNIT · StyleGAN11
⚙️ Style as Model ParametersStyle is materialized as a persistent parameter asset (weights, LoRA, learned tokens) reusable across new content.Johnson 2016 · LoRA25
🔌 Style as Learned ConditionAn external reference is encoded into a standalone style condition consumed via a dedicated conditioning interface.Ghiasi 2017 · IP-Adapter27

📌 What the taxonomy is not based on: year, backbone family, output domain, or auxiliary mechanisms (ControlNet, SDS, DDIM inversion …) — those are tags, not categories. The same perceptual objective underlies Gatys (🎯 objective), Johnson (⚙️ parameter) and Ghiasi (🔌 condition); AdaIN, StyTr² and StyleAligned are all 🧩 feature operators despite spanning CNN, Transformer and Diffusion eras.

📊 Coverage at a Glance

Carrier🖼️ Image🎬 Video🧊 3D⏳ 4D
🎯 Optimization Objective6—121
🧩 Feature Operator2221—
🌀 Latent Variable92——
⚙️ Model Parameters1861—
🔌 Learned Condition2123—

Empty cells are open problems, not omissions — e.g. ⏳ 4D stylization remains nearly untouched.

📑 Table of Contents

  1. 🎯 Style as Optimization Objective — 20 papers
  2. 🧩 Style as Feature Operator — 27 papers
  3. 🌀 Style as Latent Variable — 11 papers
  4. ⚙️ Style as Model Parameters — 25 papers
  5. 🔌 Style as Learned Condition — 27 papers
  6. 📚 Datasets
  7. 🤝 Contributing

🗂️ Papers by Style Carrier

1. 🎯 Style as Optimization Objective

Style exists only as a loss, score, or guidance signal — no transferable style representation is ever materialized.

PaperVenueYearRepLink
A neural algorithm of artistic stylearXiv2015🖼️paper
Deep photo style transferCVPR2017🖼️paper
Style transfer by relaxed optimal transport and self-similarityCVPR2019🖼️paper
Arf: Artistic radiance fieldsECCV2022🧊paper
CLIPstyler: Image style transfer with a single text conditionCVPR2022🖼️paper
SNeRF: Stylized neural implicit representations for 3D scenesSIGGRAPH Asia2022🧊paper
StyleMesh: Style transfer for indoor 3D scene reconstructionsCVPR2022🧊paper
Diffusion-based image translation using disentangled style and content representationICLR2023🖼️paper
Ref-NPR: Reference-Based Non-Photorealistic Radiance Fields for Controllable Scene StylizationCVPR2023🧊paper
3D Paintbrush: Local Stylization of 3D Shapes with Cascaded Score DistillationCVPR2024🧊paper
NeRF-Art: Text-Driven Neural Radiance Fields StylizationIEEE TVCG2024🧊paper
Balanced image stylization with style matching scoreICCV2025🖼️paper
CLIPGaussian: Universal and Multimodal Style Transfer Based on Gaussian SplattingNeurIPS2025🌐paper
GAS-NeRF: Geometry-Aware Stylization of Dynamic Radiance FieldsarXiv2025⏳paper
GT²-GS: Geometry-aware Texture Transfer for Gaussian SplattingarXiv2025🧊paper
Morpheus: Text-Driven 3D Gaussian Splat Shape and Color StylizationCVPR2025🧊paper
SGSST: Scaling Gaussian Splatting Style TransferCVPR2025🧊paper
Stylizedgs: Controllable stylization for 3d gaussian splattingTPAMI2025🧊paper
DiffStyle3D: Consistent 3D Gaussian Stylization via Attention OptimizationarXiv2026🧊paper
Fantasystyle: Controllable stylized distillation for 3d gaussian splattingAAAI2026🧊paper

2. 🧩 Style as Feature Operator

Style acts through direct matching, replacement, transport, or attention coupling of content/style activations.

PaperVenueYearRepLink
Arbitrary style transfer in real-time with adaptive instance normalizationICCV2017🖼️paper
Universal style transfer via feature transformsNeurIPS2017🖼️paper
A closed-form solution to photorealistic image stylizationECCV2018🖼️paper
Avatar-net: Multi-scale zero-shot style transfer by feature decorationCVPR2018🖼️paper
Arbitrary Style Transfer with Style-Attentional NetworksCVPR2019🖼️paper
Learning linear transformations for fast image and video style transferCVPR2019🌐paper
Photorealistic style transfer via wavelet transformsICCV2019🖼️paper
AdaAttN: Revisit attention mechanism in arbitrary neural style transferICCV2021🖼️paper
Arbitrary video style transfer via multi-channel correlationICCV2021🎬paper
Artflow: Unbiased image style transfer via reversible neural flowsCVPR2021🖼️paper
CCPL: Contrastive coherence preserving loss for versatile style transferECCV2022🌐paper
Domain enhanced arbitrary image style transfer via contrastive learningSIGGRAPH2022🖼️paper
Exact feature distribution matching for arbitrary style transfer and domain generalizationCVPR2022🖼️paper
Stytr2: Image style transfer with transformersCVPR2022🖼️paper
CAP-VSTNet: Content affinity preserved versatile style transferCVPR2023🖼️paper
Ctrl-x: Controlling structure and appearance for text-to-image generation without guidanceNeurIPS2024🖼️paper
Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion NetworkCVPR2024🖼️paper
Style aligned image generation via shared attentionCVPR2024🖼️paper
Style injection in diffusion: A training-free approach for adapting large-scale diffusion models for style transferCVPR2024🖼️paper
Training-free consistent text-to-image generationACM TOG2024🖼️paper
AttenST: A Training-Free Attention-Driven Style Transfer Framework with Pre-Trained Diffusion ModelsarXiv2025🖼️paper
DVI: Disentangling Semantic and Visual Identity for Training-Free Personalized GenerationarXiv2025🖼️paper
FPGS: Feed-Forward Semantic-aware Photorealistic Style Transfer of Large-Scale Gaussian SplattingInt. J. Comput. Vis. 20262025🧊paper
SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style TransferCVPR2025🖼️paper
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion ModelsarXiv2025🎬paper
StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style TransferCVPR2025🖼️paper
StyleStudio: Text-Driven Style Transfer with Selective Control of Style ElementsCVPR2025🖼️paper

3. 🌀 Style as Latent Variable

Style is a native generative latent with its own prior — sampleable, mixable, and editable without any reference.

PaperVenueYearRepLink
Diverse image-to-image translation via disentangled representationsECCV2018🖼️paper
MoCoGAN: Decomposing motion and content for video generationCVPR2018🎬paper
Multimodal unsupervised image-to-image translationECCV2018🖼️paper
Multimodal Unsupervised Image-to-image TranslationECCV2018🖼️paper
A style-based generator architecture for generative adversarial networksCVPR2019🖼️paper
Analyzing and Improving the Image Quality of StyleGANCVPR2020🖼️paper
Analyzing and Improving the Image Quality of StyleGANCVPR2020🖼️paper
Alias-Free Generative Adversarial NetworksNeurIPS2021🖼️paper
Styleclip: Text-driven manipulation of stylegan imageryICCV2021🖼️paper
Diffusion autoencoders: Toward a meaningful and decodable representationCVPR2022🖼️paper
Styleinv: A temporal style modulated inversion network for unconditional video generationICCV2023🎬paper

4. ⚙️ Style as Model Parameters

Style is materialized as a persistent parameter asset (weights, LoRA, learned tokens) reusable across new content.

PaperVenueYearRepLink
Perceptual losses for real-time style transfer and super-resolutionECCV2016🖼️paper
Texture networks: Feed-forward synthesis of textures and stylized imagesICML2016🖼️paper
A Learned Representation for Artistic StyleICLR2017🖼️paper
Coherent online video style transferICCV2017🎬paper
Real-time neural style transfer for videosCVPR2017🎬paper
ReCoNet: Real-time coherent video style transfer networkECCV2018🎬paper
VToonify: Controllable High-Resolution Portrait Video Style TransferACM TOG2022🎬paper
An image is worth one word: Personalizing text-to-image generation using textual inversionICLR2023🖼️paper
Break-a-scene: Extracting multiple concepts from a single imageSIGGRAPH Asia2023🖼️paper
Cones: Concept neurons in diffusion models for customized generationICML2023🖼️paper
DreamBooth3D: Subject-Driven Text-to-3D GenerationICCV2023🧊paper
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generationCVPR2023🖼️paper
Inversion-Based Style Transfer With Diffusion ModelsCVPR2023🖼️paper
Key-locked rank one editing for text-to-image personalizationSIGGRAPH2023🖼️paper
Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion modelsNeurIPS2023🖼️paper
Multi-concept customization of text-to-image diffusionCVPR2023🖼️paper
Prospect: Prompt spectrum for attribute-aware personalization of diffusion modelsACM TOG2023🖼️paper
Styledrop: Text-to-image generation in any stylearXiv2023🖼️paper
Dreamvideo: Composing your dream videos with customized subject and motionCVPR2024🎬paper
Implicit style-content separation using b-loraECCV2024🖼️paper
Magic-me: Identity-specific video customized diffusionECCV2024🎬paper
ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAsECCV2024🖼️paper
APT: Adaptive Personalized Training for Diffusion Models with Limited DataCVPR2025🖼️paper
ConsisLoRA: Enhancing Content and Style Consistency for LoRA-based Style TransferarXiv2025🖼️paper
Photodoodle: Learning artistic image editing from few-shot pairwise dataarXiv2025🖼️paper

5. 🔌 Style as Learned Condition

An external reference is encoded into a standalone style condition consumed via a dedicated conditioning interface.

PaperVenueYearRepLink
CLIP-NeRF: Text-and-image driven manipulation of neural radiance fieldsCVPR2022🧊paper
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editingNeurIPS2023🖼️paper
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generationICCV2023🖼️paper
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion ModelsarXiv2023🖼️paper
Arc2Face: A Foundation Model for ID-Consistent Human FacesECCV2024🖼️paper
ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationCVPR2024🖼️paper
Deadiff: An efficient stylization diffusion model with disentangled representationsCVPR2024🖼️paper
DiffStyler: Diffusion-based Localized Image Style TransferarXiv2024🖼️paper
Instantid: Zero-shot identity-preserving generation in secondsarXiv2024🖼️paper
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image GenerationarXiv2024🖼️paper
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image GenerationarXiv2024🖼️paper
Photomaker: Customizing realistic human photos via stacked id embeddingCVPR2024🖼️paper
Pulid: Pure and lightning id customization via contrastive alignmentNeurIPS2024🖼️paper
Ssr-encoder: Encoding selective subject representation for subject-driven generationCVPR2024🖼️paper
StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter LearningACM TOG2024🎬paper
CSGO: Content-Style Composition in Text-to-Image GenerationNeurIPS2025🖼️paper
Free-lunch color-texture disentanglement for stylized image generationNeurIPS2025🖼️paper
ID-Consistent, Precise Expression Generation with Blendshape-Guided DiffusionICCV2025🖼️paper
Nested attention: Semantic-aware attention values for concept personalizationSIGGRAPH2025🖼️paper
Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationarXiv2025🖼️paper
Stylemaster: Stylize your video with artistic generation and translationCVPR2025🎬paper
Styleshot: A snapshot on any styleTPAMI2025🖼️paper
Stylos: Multi-View 3D Stylization with Single-Forward Gaussian SplattingarXiv2025🧊paper
U-StyDiT: Ultra-high quality artistic style transfer using diffusion transformersarXiv preprint arXiv:2503.081572025🖼️paper
Uso: Unified style and subject-driven generation via disentangled and reward learningarXiv preprint arXiv:2508.189662025🖼️paper
AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian SplattingarXiv2026🧊paper
TeleStyle: Content-Preserving Style Transfer in Images and VideosarXiv preprint arXiv:2601.201752026🌐paper

📚 Datasets

🖼️ Image Datasets
DatasetYearSizeDescriptionLink
WikiArt201842,129 imagesStyle, Artist, Genrelink
Stylized ImageNet2018~134GBStyle Transferlink
BAM20172.5M imagesArtistic Medialink
ArtBench-10202260,000 imagesArtwork Benchmarklink
FFHQ201970,000 imagesHuman Faceslink
MetFaces20201,336 imagesArtistic Faceslink
AAHQ202125,000 imagesArtistic Faceslink
Ukiyo-e Faces20205,209 imagesAligned Ukiyo-e Faceslink
DiffusionDB202214M imagesText-to-imagelink
JourneyDB20234.4M imagesMultimodal Visionlink
StyleShot2024-Style Transferlink
Danbooru201720172.94M imagesAnimelink
Chinese Style Transfer20181000 content, 100 styleChinese Paintinglink
4SKST202325 color, 100 sketchesSketch Stylelink
🎬 Video Datasets
DatasetYearSizeDescriptionLink
UADFV2018100 videosVideo Style Transferlink
FaceForensics++20196000 videosSwapped Facelink
Celeb-DF2020408 original videosDeepFakelink
DFDC2020100,000 clipsDeepFake Detectionlink
FFIW-10K202110,000 videosFace Forgerylink
ForgeryNet2021221,247 videosForgery Analysislink
🧊 3D / Motion Datasets
DatasetYearSizeDescriptionLink
100STYLE20224M framesStylized Motion Capturelink
Motiondataset202336,673 frames3D Motionlink

🤝 Contributing

We welcome contributions! When suggesting a paper, please indicate which style carrier it belongs to:

  • 🎯 Objective — style lives only in the loss / score / guidance
  • 🧩 Feature — direct activation matching, transport, or attention coupling
  • 🌀 Latent — native generative latent, sampleable without a reference
  • ⚙️ Parameter — persistent style asset (weights / LoRA / tokens)
  • 🔌 Condition — reference-encoded standalone style condition

Ways to help:

  • 🐛 Report bugs and issues
  • 💡 Suggest new papers or resources
  • 🔧 Submit pull requests
  • ⭐ Star this repository if you find it helpful!

📖 Citation

If you find this repository useful for your research, please consider citing:

@article{zhang2025style,
  title={Style transfer: A decade survey},
  author={Zhang, Tianshan and Tang, Hao},
  journal={arXiv preprint arXiv:2506.19278},
  year={2025}
}

⭐ Star History

Star History Chart

🙏 Acknowledgments

Thanks to all researchers and developers who made their work publicly available.

By Monet's Impression of Sunrise

Contributors

neptune-T

79 commits

00lostin00

2 commits