mlpc-lab/TokenCompose_SD21_B

Model

🧩 TokenCompose SD21 Model Card

1

8 commits

3 linked in READMEs

updated Jul 6, 2024

See the code

README

🧩 TokenCompose SD21 Model Card

🎬CVPR 2024

TokenCompose_SD21_B is a latent text-to-image diffusion model finetuned from the Stable-Diffusion-v2-1 checkpoint at resolution 768x768 on the VSR split of COCO image-caption pairs for 32,000 steps with a learning rate of 5e-6. The training objective involves token-level grounding terms in addition to denoising loss for enhanced multi-category instance composition and photorealism. The "_A/B" postfix indicates different finetuning runs of the model using the same above configurations.

📄 Paper

Please follow this link.

🧨Example Usage

We strongly recommend using the 🤗Diffuser library to run our model.

import torch
from diffusers import StableDiffusionPipeline

model_id = "mlpc-lab/TokenCompose_SD21_B"
device = "cuda"

pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float32)
pipe = pipe.to(device)

prompt = "A cat and a wine glass"
image = pipe(prompt).images[0]  
    
image.save("cat_and_wine_glass.png")

⬆️Improvements over SD21

ModelObject AccuracyMG3 COCOMG4 COCOMG5 COCOMG3 ADE20KMG4 ADE20KMG5 ADE20KFID COCO
SD2147.8270.1425.573.2775.1335.077.1619.59
TokenCompose (SD21)60.1080.4836.695.7179.5139.598.1319.15

📰 Citation

@InProceedings{Wang2024TokenCompose,
    author    = {Wang, Zirui and Sha, Zhizhou and Ding, Zheng and Wang, Yilin and Tu, Zhuowen},
    title     = {TokenCompose: Text-to-Image Diffusion with Token-level Supervision},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2024},
    pages     = {8553-8564}
}
compositionality
cvpr
diffusers
endpoints_compatible
image-generation
safetensors
stable-diffusion
text-to-image

mlpc-lab/TokenCompose_SD21_B

Model

🧩 TokenCompose SD21 Model Card

1

8 commits

3 linked in READMEs

updated Jul 6, 2024

See the code

README

🧩 TokenCompose SD21 Model Card

🎬CVPR 2024

TokenCompose_SD21_B is a latent text-to-image diffusion model finetuned from the Stable-Diffusion-v2-1 checkpoint at resolution 768x768 on the VSR split of COCO image-caption pairs for 32,000 steps with a learning rate of 5e-6. The training objective involves token-level grounding terms in addition to denoising loss for enhanced multi-category instance composition and photorealism. The "_A/B" postfix indicates different finetuning runs of the model using the same above configurations.

📄 Paper

Please follow this link.

🧨Example Usage

We strongly recommend using the 🤗Diffuser library to run our model.

import torch
from diffusers import StableDiffusionPipeline

model_id = "mlpc-lab/TokenCompose_SD21_B"
device = "cuda"

pipe = StableDiffusionPipeline.from_pretrained(model_id, torch_dtype=torch.float32)
pipe = pipe.to(device)

prompt = "A cat and a wine glass"
image = pipe(prompt).images[0]  
    
image.save("cat_and_wine_glass.png")

⬆️Improvements over SD21

ModelObject AccuracyMG3 COCOMG4 COCOMG5 COCOMG3 ADE20KMG4 ADE20KMG5 ADE20KFID COCO
SD2147.8270.1425.573.2775.1335.077.1619.59
TokenCompose (SD21)60.1080.4836.695.7179.5139.598.1319.15

📰 Citation

@InProceedings{Wang2024TokenCompose,
    author    = {Wang, Zirui and Sha, Zhizhou and Ding, Zheng and Wang, Yilin and Tu, Zhuowen},
    title     = {TokenCompose: Text-to-Image Diffusion with Token-level Supervision},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
    month     = {June},
    year      = {2024},
    pages     = {8553-8564}
}
compositionality
cvpr
diffusers
endpoints_compatible
image-generation
safetensors
stable-diffusion
text-to-image