mlpc-lab/TokenCompose_SD14_A

Model

🧩 TokenCompose SD14 Model Card

2

13 commits

3 linked in READMEs

updated Jul 6, 2024

See the code

README

🧩 TokenCompose SD14 Model Card

🎬CVPR 2024

TokenCompose_SD14_A is a latent text-to-image diffusion model finetuned from the Stable-Diffusion-v1-4 checkpoint at resolution 512x512 on the VSR split of COCO image-caption pairs for 24,000 steps with a learning rate of 5e-6. The training objective involves token-level grounding terms in addition to denoising loss for enhanced multi-category instance composition and photorealism. The "_A/B" postfix indicates different finetuning runs of the model using the same above configurations.

📄 Paper

Please follow this link.

🧨Example Usage

We strongly recommend using the 🤗Diffuser library to run our model.

import torch
from diffusers import StableDiffusionPipeline

model_id = "mlpc-lab/TokenCompose_SD14_A"
device = "cuda"

pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float32)
pipe = pipe.to(device)

prompt = "A cat and a wine glass"
image = pipe(prompt).images[0]  
    
image.save("cat_and_wine_glass.png")

⬆️Improvements over SD14

MethodMulti-category Instance CompositionPhotorealismEfficiency
Object AccuracyCOCOADE20KFID (COCO)FID (Flickr30K)Latency
MG2MG3MG4MG5MG2MG3MG4MG5
SD 1.429.8690.721.3350.740.8911.680.450.880.2189.810.4053.961.1416.521.131.890.3420.8871.467.540.17
TokenCompose (Ours)52.1598.080.4076.161.0428.810.953.280.4897.750.3476.931.0933.921.476.210.6220.1971.137.560.14

📰 Citation

@InProceedings{Wang2024TokenCompose,
    author    = {Wang, Zirui and Sha, Zhizhou and Ding, Zheng and Wang, Yilin and Tu, Zhuowen},
    title     = {TokenCompose: Text-to-Image Diffusion with Token-level Supervision},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2024},
    pages     = {8553-8564}
}
compositionality
cvpr
diffusers
endpoints_compatible
image-generation
safetensors
stable-diffusion
text-to-image

mlpc-lab/TokenCompose_SD14_A

Model

🧩 TokenCompose SD14 Model Card

2

13 commits

3 linked in READMEs

updated Jul 6, 2024

See the code

README

🧩 TokenCompose SD14 Model Card

🎬CVPR 2024

TokenCompose_SD14_A is a latent text-to-image diffusion model finetuned from the Stable-Diffusion-v1-4 checkpoint at resolution 512x512 on the VSR split of COCO image-caption pairs for 24,000 steps with a learning rate of 5e-6. The training objective involves token-level grounding terms in addition to denoising loss for enhanced multi-category instance composition and photorealism. The "_A/B" postfix indicates different finetuning runs of the model using the same above configurations.

📄 Paper

Please follow this link.

🧨Example Usage

We strongly recommend using the 🤗Diffuser library to run our model.

import torch
from diffusers import StableDiffusionPipeline

model_id = "mlpc-lab/TokenCompose_SD14_A"
device = "cuda"

pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float32)
pipe = pipe.to(device)

prompt = "A cat and a wine glass"
image = pipe(prompt).images[0]  
    
image.save("cat_and_wine_glass.png")

⬆️Improvements over SD14

MethodMulti-category Instance CompositionPhotorealismEfficiency
Object AccuracyCOCOADE20KFID (COCO)FID (Flickr30K)Latency
MG2MG3MG4MG5MG2MG3MG4MG5
SD 1.429.8690.721.3350.740.8911.680.450.880.2189.810.4053.961.1416.521.131.890.3420.8871.467.540.17
TokenCompose (Ours)52.1598.080.4076.161.0428.810.953.280.4897.750.3476.931.0933.921.476.210.6220.1971.137.560.14

📰 Citation

@InProceedings{Wang2024TokenCompose,
    author    = {Wang, Zirui and Sha, Zhizhou and Ding, Zheng and Wang, Yilin and Tu, Zhuowen},
    title     = {TokenCompose: Text-to-Image Diffusion with Token-level Supervision},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2024},
    pages     = {8553-8564}
}
compositionality
cvpr
diffusers
endpoints_compatible
image-generation
safetensors
stable-diffusion
text-to-image