A curated list of awesome research papers, projects, code, dataset, workshops etc. related to virtual try-off.
56
37 commits
updated Sep 5, 2026
|
Virtual Try-On (VTON) takes a Masked Ref. Image $\mathrm{I}'$ and a Garment $\mathrm{G}$ and synthesizes a dressed person, the Predicted $\hat{\mathrm{I}}$. Virtual Try-Off (VTOFF) is the inverse: from a clothed Reference $\mathrm{I}$, it reconstructs a canonical Predicted Garment $\hat{\mathrm{G}}$. Each prediction is scored against its ground truth with image similarity metrics: VTON Loss compares $\hat{\mathrm{I}}$ to $\mathrm{I}$, and VTOFF Loss compares $\hat{\mathrm{G}}$ to $\mathrm{G}$. The two tasks also form a cycle: the output of one can serve as the input to the other. VTOFF was introduced and the term coined in TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models. Figure from TryOffDiff. |
TryOffDiff.FLUX.1-dev and the LoRA "prithivMLmods/Canopus-Clothing-Flux-LoRA" to generate clothes from text.FLUX.1-dev and CatVTON.FLUX.1-dev-Redux + FLUX.1-dev-Depth. Set mask as 'structure' and model image as 'style' for targeting VTOFF.FLUX.2-klein-9B for vtoff task. They also generate 360-video of the garment with a workflow.[2025-06-17]
FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
[paper]
experimented with task, see Figure 6.
[2026-03-10]
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
[paper]
experimented with task, see Figure 16.
[2026-04-09]
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
[paper]
used Nano Banana Pro as VTOFF method to generate paired dataset.
[2026-06-11]
A VLM-based framework for evaluating garment consistency in AI-generated images based on DeLong’s theory
[paper]
developed a garment-specific evaluation framework, surpassing traditional eval metrics.
[2020-05-08]
TileGAN: category-oriented attention-based high-quality tiled clothes generation from dressed person
[paper]
targeted vtoff task with a two-stage method.
[2021-08-08]
ViTon-GUN: Person-to-Person Virtual Try-on via Garment Unwrapping
[paper]
targeted p2p-vton task: run vtoff first, then vton.
[2022-08-11]
ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design
[paper]
text-to-garment generation.
[2023-08-15]
SGDiff: A Style Guided Diffusion Model for Fashion Synthesis
[paper]
finetuned GLIDE to generate garment images from "text" (garment attributes) + "style image".
[2023-08-22]
DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment
[paper]
finetuned StableDiffusion for targeting garment synthesis and manipulation.
[2024-01-29]
DressCode: Autoregressively Sewing and Generating Garments from Text Guidance
[paper]
[project]
[code]
text-to-3dGarment with 2 branches: text-to-sewing-patterns with SewingGPT (3d garment) and text-to-texture with finetuned StableDiffusion.
[2024-04-22]
FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on
[paper]
[code]
vtoff task is included in the loss function for training a vton model, vtoff was not the focus of the paper, nor was a stand-alone task introduced.
[2024-04-26]
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
[paper]
finetuned SD-v1.5 with ControlNet to generate garments from text + sketch image.
[2024-11-25]
Controllable Human Image Generation with Personalized Multi-Garments
[paper]
[project]
[code]
generated garment images from segmented garments—extracted from reference images—to construct a dataset for training a multi-garment VTON model.
[2024-11-27]
TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models
[paper]
[project]
[code]
coined the term Virtual Try-Off (VTOFF), formally introduced the task and introduced the baselines.
[2024-11-29]
RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation
[paper]
[2024-12-11]
TryOffAnyone: Tiled Cloth Generation from a Dressed Person
[paper]
[code]
finetuned Stable Diffusion-v1.5-inpainting for vtoff in CatVTON-style.
[2024-12-16]
IGR: Improving Diffusion Model for Garment Restoration from Person Image
[paper]
finetuned Stable Diffusion-v1.5 for vtoff using 2 UNets (Reference and Denoiser).
[2025-01-08]
Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling
[paper]
used VTOFF to generate synthetic pairs to enrich VTON training dataset. Trained a U-Net model for warping latent features.
[2025-01-27]
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
[paper]
[project]
[code]
1st unified model targeting both VTON and VTOFF.
[2025-03-07]
DiffDesign: A diffusion model using garment Knowledge-Enhanced for Fashion Design Synthesis
[paper]
trained a WURSTCHEN-based model for text-to-garment task.
[2025-04-17]
MGT: Extending Virtual Try-Off to Multi-Garment Scenarios
[paper]
[project]
[code]
1st VTOFF model supporting multiple garment reconstruction. Finetuned StableDiffusion-v1.4 (similar to TryOffDiff) and trained a class embedding for multi-garment support.
[2025-04-17]
IMAGGarment-1: Fine-Grained Garment Generation for Controllable Fashion Design
[paper]
[project]
[code]
garment synthesis with precise control over silhouette, color, and logo placement (3 inputs). Incorporates a tower architecure with SD-1.5 for coarse generation followed by a fine-grained approach utilizing SD-1.5-inpaint.
[2025-05-26]
ImgEdit: A Unified Image Editing Dataset and Benchmark
[paper]
[code]
trained an image editing model using Vision-Language Model to process a reference image and editing prompt, targeting multiple tasks including VTOFF.
[2025-05-27]
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
[paper]
[project]
[code]
1st dual-DiT model for VTOFF with multi-garment support, built on a finetuned StableDiffusion-v3 with modified attention. It accepts images, text, or masks as input, enabling multi-category garment handling.
[2025-08-06]
Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
[paper]
unified framework (TWGTM) targeting both VTON & VTOFF.
[2025-08-06]
One Model For All: Partial Diffusion for Unified Try-On and Try-Off in Any Pose
[paper]
[project]
unified framework (OMFA) targeting both VTON & VTOFF.
[2025-08-06]
Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
[paper]
[project]
unified framework (Voost) targeting both VTON & VTOFF.
[2025-08-25]
JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
[paper]
trained a VTOFF model for cyclic data generation pipeline.
[2025-11-19]
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
[paper]
[code]
a universal framework driven by Multimodal Large Language Model.
[2026-01-05]
AlignVTOFF: Texture-Spatial Feature Alignment for High-Fidelity Virtual Try-Off
[paper]
dual-U-Net pipeline trained with a loss combining regular diffusion loss + LPIPS.
[2026-03-10]
BridgeDiff: Bridging Human Observations and Flat-Garment Synthesis for Virtual Try-Off
[paper]
dual-U-Net framework based on Stable-Diffusion-v1.5.
[2026-03-23]
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
[paper]
introduced a synthetic benchmark Dress Editing Dataset (Dress-ED) and Dress Editing Model (Dress-EM) targeting both VTON and VTOFF.
[2026-03-24]
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
[paper]
introduced a unified Omni-VTON framework based on the Diffusion Transformer (DiT) that combines model-based VTON, model-free VTON, and VTOFF tasks in a single mask-free model, trained on a self-curated Omni-TryOn dataset of over 380k image pairs.
[2026-04-09]
What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction
[paper]
introduced a Dual-UNet Diffusion Model for VTOFF, thoroughly ablating design choices in generation backbones, conditioning (masks, inputs, semantics), and training strategies/losses to reconstruct canonical garments from draped images.
[2026-07-09]
MMTryOff: multi-category virtual try-off with mask-free inference via diffusion transformer
[paper]
LoRA training with FLUX.1-dev, on newly proposed VITOFF-HD dataset, incorporating frequency loss and mask loss during training. CatVTON-style training. No code, no dataset available.
[2026-08-29]
RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation
[paper]
follow-up to RAGDiffusion: Dual-Image-Stream FLUX + AR-GRPO RL post-training with Garment-RM to recover high-frequency textures (weaves, logos) after RAG’s macro-constraints. Introduces STGarment-Plus.
A curated list of awesome research papers, projects, code, dataset, workshops etc. related to virtual try-off.
56
37 commits
updated Sep 5, 2026
|
Virtual Try-On (VTON) takes a Masked Ref. Image $\mathrm{I}'$ and a Garment $\mathrm{G}$ and synthesizes a dressed person, the Predicted $\hat{\mathrm{I}}$. Virtual Try-Off (VTOFF) is the inverse: from a clothed Reference $\mathrm{I}$, it reconstructs a canonical Predicted Garment $\hat{\mathrm{G}}$. Each prediction is scored against its ground truth with image similarity metrics: VTON Loss compares $\hat{\mathrm{I}}$ to $\mathrm{I}$, and VTOFF Loss compares $\hat{\mathrm{G}}$ to $\mathrm{G}$. The two tasks also form a cycle: the output of one can serve as the input to the other. VTOFF was introduced and the term coined in TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models. Figure from TryOffDiff. |
TryOffDiff.FLUX.1-dev and the LoRA "prithivMLmods/Canopus-Clothing-Flux-LoRA" to generate clothes from text.FLUX.1-dev and CatVTON.FLUX.1-dev-Redux + FLUX.1-dev-Depth. Set mask as 'structure' and model image as 'style' for targeting VTOFF.FLUX.2-klein-9B for vtoff task. They also generate 360-video of the garment with a workflow.[2025-06-17]
FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
[paper]
experimented with task, see Figure 6.
[2026-03-10]
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
[paper]
experimented with task, see Figure 16.
[2026-04-09]
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
[paper]
used Nano Banana Pro as VTOFF method to generate paired dataset.
[2026-06-11]
A VLM-based framework for evaluating garment consistency in AI-generated images based on DeLong’s theory
[paper]
developed a garment-specific evaluation framework, surpassing traditional eval metrics.
[2020-05-08]
TileGAN: category-oriented attention-based high-quality tiled clothes generation from dressed person
[paper]
targeted vtoff task with a two-stage method.
[2021-08-08]
ViTon-GUN: Person-to-Person Virtual Try-on via Garment Unwrapping
[paper]
targeted p2p-vton task: run vtoff first, then vton.
[2022-08-11]
ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design
[paper]
text-to-garment generation.
[2023-08-15]
SGDiff: A Style Guided Diffusion Model for Fashion Synthesis
[paper]
finetuned GLIDE to generate garment images from "text" (garment attributes) + "style image".
[2023-08-22]
DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment
[paper]
finetuned StableDiffusion for targeting garment synthesis and manipulation.
[2024-01-29]
DressCode: Autoregressively Sewing and Generating Garments from Text Guidance
[paper]
[project]
[code]
text-to-3dGarment with 2 branches: text-to-sewing-patterns with SewingGPT (3d garment) and text-to-texture with finetuned StableDiffusion.
[2024-04-22]
FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on
[paper]
[code]
vtoff task is included in the loss function for training a vton model, vtoff was not the focus of the paper, nor was a stand-alone task introduced.
[2024-04-26]
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
[paper]
finetuned SD-v1.5 with ControlNet to generate garments from text + sketch image.
[2024-11-25]
Controllable Human Image Generation with Personalized Multi-Garments
[paper]
[project]
[code]
generated garment images from segmented garments—extracted from reference images—to construct a dataset for training a multi-garment VTON model.
[2024-11-27]
TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models
[paper]
[project]
[code]
coined the term Virtual Try-Off (VTOFF), formally introduced the task and introduced the baselines.
[2024-11-29]
RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation
[paper]
[2024-12-11]
TryOffAnyone: Tiled Cloth Generation from a Dressed Person
[paper]
[code]
finetuned Stable Diffusion-v1.5-inpainting for vtoff in CatVTON-style.
[2024-12-16]
IGR: Improving Diffusion Model for Garment Restoration from Person Image
[paper]
finetuned Stable Diffusion-v1.5 for vtoff using 2 UNets (Reference and Denoiser).
[2025-01-08]
Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling
[paper]
used VTOFF to generate synthetic pairs to enrich VTON training dataset. Trained a U-Net model for warping latent features.
[2025-01-27]
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
[paper]
[project]
[code]
1st unified model targeting both VTON and VTOFF.
[2025-03-07]
DiffDesign: A diffusion model using garment Knowledge-Enhanced for Fashion Design Synthesis
[paper]
trained a WURSTCHEN-based model for text-to-garment task.
[2025-04-17]
MGT: Extending Virtual Try-Off to Multi-Garment Scenarios
[paper]
[project]
[code]
1st VTOFF model supporting multiple garment reconstruction. Finetuned StableDiffusion-v1.4 (similar to TryOffDiff) and trained a class embedding for multi-garment support.
[2025-04-17]
IMAGGarment-1: Fine-Grained Garment Generation for Controllable Fashion Design
[paper]
[project]
[code]
garment synthesis with precise control over silhouette, color, and logo placement (3 inputs). Incorporates a tower architecure with SD-1.5 for coarse generation followed by a fine-grained approach utilizing SD-1.5-inpaint.
[2025-05-26]
ImgEdit: A Unified Image Editing Dataset and Benchmark
[paper]
[code]
trained an image editing model using Vision-Language Model to process a reference image and editing prompt, targeting multiple tasks including VTOFF.
[2025-05-27]
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
[paper]
[project]
[code]
1st dual-DiT model for VTOFF with multi-garment support, built on a finetuned StableDiffusion-v3 with modified attention. It accepts images, text, or masks as input, enabling multi-category garment handling.
[2025-08-06]
Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
[paper]
unified framework (TWGTM) targeting both VTON & VTOFF.
[2025-08-06]
One Model For All: Partial Diffusion for Unified Try-On and Try-Off in Any Pose
[paper]
[project]
unified framework (OMFA) targeting both VTON & VTOFF.
[2025-08-06]
Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
[paper]
[project]
unified framework (Voost) targeting both VTON & VTOFF.
[2025-08-25]
JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
[paper]
trained a VTOFF model for cyclic data generation pipeline.
[2025-11-19]
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
[paper]
[code]
a universal framework driven by Multimodal Large Language Model.
[2026-01-05]
AlignVTOFF: Texture-Spatial Feature Alignment for High-Fidelity Virtual Try-Off
[paper]
dual-U-Net pipeline trained with a loss combining regular diffusion loss + LPIPS.
[2026-03-10]
BridgeDiff: Bridging Human Observations and Flat-Garment Synthesis for Virtual Try-Off
[paper]
dual-U-Net framework based on Stable-Diffusion-v1.5.
[2026-03-23]
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
[paper]
introduced a synthetic benchmark Dress Editing Dataset (Dress-ED) and Dress Editing Model (Dress-EM) targeting both VTON and VTOFF.
[2026-03-24]
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
[paper]
introduced a unified Omni-VTON framework based on the Diffusion Transformer (DiT) that combines model-based VTON, model-free VTON, and VTOFF tasks in a single mask-free model, trained on a self-curated Omni-TryOn dataset of over 380k image pairs.
[2026-04-09]
What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction
[paper]
introduced a Dual-UNet Diffusion Model for VTOFF, thoroughly ablating design choices in generation backbones, conditioning (masks, inputs, semantics), and training strategies/losses to reconstruct canonical garments from draped images.
[2026-07-09]
MMTryOff: multi-category virtual try-off with mask-free inference via diffusion transformer
[paper]
LoRA training with FLUX.1-dev, on newly proposed VITOFF-HD dataset, incorporating frequency loss and mask loss during training. CatVTON-style training. No code, no dataset available.
[2026-08-29]
RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation
[paper]
follow-up to RAGDiffusion: Dual-Image-Stream FLUX + AR-GRPO RL post-training with Garment-RM to recover high-frequency textures (weaves, logos) after RAG’s macro-constraints. Introduces STGarment-Plus.