Training DALLE from scratch, utilizing target language's PLMs' token embedding layer and position embedding layer as text encoder.
π For the project details, please refer to README.pdf
| OpenAIβs DALLE | KoDALLE of HappyFace | |
|---|---|---|
| Train Dataset Size | 250 Million Pairs | 0.8 Million Pairs |
| #Params | 12 Billion | 428 Million |
| #Layers | 64 Layers | 16 Layers |
| Computing Resource | 1024 x V100 16GB | 1 x V100 32GB |
| Text Encoder | 16384 Vocab x 512 Dim BPE | 32000 Vocab x 1024 Dim klue/roberta-large |
| Image Encoder | VQVAE | VQGAN |
| Optimizer | AdamW | AdamW |
| Learning Rate | 4.5e-5 | 3.0e-5 |
| Weight Decay | 4.5e-3 | 3.0e-3 |
| LR Scheduler | ReduceLROnPlateau | - |
The team constructed Text to Fashion Design DALLE model in Korean language with less than 100k text-image sampled pairs.
| Caption | νμμμ μμμ μ€μΉ΄μ΄λΈλ£¨μ΄λ€. μμμμ κΈ°μ₯μ λ‘±μ΄λ€. μμμ νμ΄νΈμ΄λ€. μΉ΄ν κ³ λ¦¬λ λΈλΌμ°μ€μ΄λ€. λν μΌμλ μ λ§μ΄λ€. μλ§€κΈ°μ₯μ λ°νμ΄λ€. μμ¬μλ μ€ν¬μ΄λ€. νλ¦°νΈμλ 무μ§μ΄λ€. λ₯λΌμΈμ λΈμ΄λ₯μ΄λ€. νμ λ Έλ© |
| Generated Image | ![]() |
| Caption | μμ°ν°λ μμμ΄ μΉ΄ν€ μμ¬κ° μ°λΈ νμ΄ λ£¨μ¦μΈ μ½νΈμ΄λ€. νμλ μμμ΄ λ€μ΄λΉ μμ¬κ° λ°λ νμ΄ μ€ν€λμΈ μ²λ°μ§μ΄λ€. |
| Generated Image | ![]() |
| Caption | νμμμ κΈ°μ₯μ λ°λͺ©μ΄λ€. μμμ λΈλ£¨μ΄λ€. μΉ΄ν κ³ λ¦¬λ μ€μ»€νΈμ΄λ€. μμ¬μλ λ°λμ΄λ€. νμ μμ΄λμ΄λ€. μμμμ μμμ νμ΄νΈμ΄λ€. μΉ΄ν κ³ λ¦¬λ λΈλΌμ°μ€μ΄λ€. λν μΌμλ μ λ§μ΄λ€. μλ§€κΈ°μ₯μ λ°νμ΄λ€. μμ¬μλ μ°λΈμ΄λ€. |
| Generated Image | ![]() |
| Caption | μμμμ κΈ°μ₯μ λ Έλ©μ΄λ€. μμμμ μμμ νμ΄νΈμ΄λ€. μμμμ μλΈμμμ λΈλμ΄λ€. μμμμ μΉ΄ν κ³ λ¦¬λ ν°μ μΈ μ΄λ€. μμμμ μλ§€κΈ°μ₯μ λ°νμ΄λ€. μμμμ μμ¬μλ μ μ§μ΄λ€. μμμμ νλ¦°νΈμλ λ ν°λ§μ΄λ€. μμμμ λ₯λΌμΈμ λΌμ΄λλ₯μ΄λ€. μμμμ νμ 루μ¦μ΄λ€. |
| Generated Image | ![]() |
Experimentations were conducted with the following Korean Transformers Modelsβ embedding layers. The team selected klue/roberta-large as baseline in the repository considering the size of the model.
KoDALLE with klue/roberta-large's wpe and wte were trained on 32GB V100 GPU environment. Hyperparams related to the DALLE's model size are following.
'BATCH_SIZE': 40
'DEPTH': 16
'TEXT_SEQ_LEN': 128
'VOCAB_SIZE': 32000
'MODEL_DIM': 1024
'ATTN_TYPES': 'full'
'DIM_HEAD': 64
'HEADS': 8
@misc{ramesh2021zeroshot,
title = {Zero-Shot Text-to-Image Generation},
author = {Aditya Ramesh and Mikhail Pavlov and Gabriel Goh and Scott Gray and Chelsea Voss and Alec Radford and Mark Chen and Ilya Sutskever},
year = {2021},
eprint = {2102.12092},
archivePrefix = {arXiv},
primaryClass = {cs.CV}
}
@misc{esser2021taming,
title = {Taming Transformers for High-Resolution Image Synthesis},
author = {Patrick Esser and Robin Rombach and BjΓΆrn Ommer},
year = {2021},
eprint = {2012.09841},
archivePrefix = {arXiv},
primaryClass = {cs.CV}
}
Python
100.0%
Training DALLE from scratch, utilizing target language's PLMs' token embedding layer and position embedding layer as text encoder.
π For the project details, please refer to README.pdf
| OpenAIβs DALLE | KoDALLE of HappyFace | |
|---|---|---|
| Train Dataset Size | 250 Million Pairs | 0.8 Million Pairs |
| #Params | 12 Billion | 428 Million |
| #Layers | 64 Layers | 16 Layers |
| Computing Resource | 1024 x V100 16GB | 1 x V100 32GB |
| Text Encoder | 16384 Vocab x 512 Dim BPE | 32000 Vocab x 1024 Dim klue/roberta-large |
| Image Encoder | VQVAE | VQGAN |
| Optimizer | AdamW | AdamW |
| Learning Rate | 4.5e-5 | 3.0e-5 |
| Weight Decay | 4.5e-3 | 3.0e-3 |
| LR Scheduler | ReduceLROnPlateau | - |
The team constructed Text to Fashion Design DALLE model in Korean language with less than 100k text-image sampled pairs.
| Caption | νμμμ μμμ μ€μΉ΄μ΄λΈλ£¨μ΄λ€. μμμμ κΈ°μ₯μ λ‘±μ΄λ€. μμμ νμ΄νΈμ΄λ€. μΉ΄ν κ³ λ¦¬λ λΈλΌμ°μ€μ΄λ€. λν μΌμλ μ λ§μ΄λ€. μλ§€κΈ°μ₯μ λ°νμ΄λ€. μμ¬μλ μ€ν¬μ΄λ€. νλ¦°νΈμλ 무μ§μ΄λ€. λ₯λΌμΈμ λΈμ΄λ₯μ΄λ€. νμ λ Έλ© |
| Generated Image | ![]() |
| Caption | μμ°ν°λ μμμ΄ μΉ΄ν€ μμ¬κ° μ°λΈ νμ΄ λ£¨μ¦μΈ μ½νΈμ΄λ€. νμλ μμμ΄ λ€μ΄λΉ μμ¬κ° λ°λ νμ΄ μ€ν€λμΈ μ²λ°μ§μ΄λ€. |
| Generated Image | ![]() |
| Caption | νμμμ κΈ°μ₯μ λ°λͺ©μ΄λ€. μμμ λΈλ£¨μ΄λ€. μΉ΄ν κ³ λ¦¬λ μ€μ»€νΈμ΄λ€. μμ¬μλ λ°λμ΄λ€. νμ μμ΄λμ΄λ€. μμμμ μμμ νμ΄νΈμ΄λ€. μΉ΄ν κ³ λ¦¬λ λΈλΌμ°μ€μ΄λ€. λν μΌμλ μ λ§μ΄λ€. μλ§€κΈ°μ₯μ λ°νμ΄λ€. μμ¬μλ μ°λΈμ΄λ€. |
| Generated Image | ![]() |
| Caption | μμμμ κΈ°μ₯μ λ Έλ©μ΄λ€. μμμμ μμμ νμ΄νΈμ΄λ€. μμμμ μλΈμμμ λΈλμ΄λ€. μμμμ μΉ΄ν κ³ λ¦¬λ ν°μ μΈ μ΄λ€. μμμμ μλ§€κΈ°μ₯μ λ°νμ΄λ€. μμμμ μμ¬μλ μ μ§μ΄λ€. μμμμ νλ¦°νΈμλ λ ν°λ§μ΄λ€. μμμμ λ₯λΌμΈμ λΌμ΄λλ₯μ΄λ€. μμμμ νμ 루μ¦μ΄λ€. |
| Generated Image | ![]() |
Experimentations were conducted with the following Korean Transformers Modelsβ embedding layers. The team selected klue/roberta-large as baseline in the repository considering the size of the model.
KoDALLE with klue/roberta-large's wpe and wte were trained on 32GB V100 GPU environment. Hyperparams related to the DALLE's model size are following.
'BATCH_SIZE': 40
'DEPTH': 16
'TEXT_SEQ_LEN': 128
'VOCAB_SIZE': 32000
'MODEL_DIM': 1024
'ATTN_TYPES': 'full'
'DIM_HEAD': 64
'HEADS': 8
@misc{ramesh2021zeroshot,
title = {Zero-Shot Text-to-Image Generation},
author = {Aditya Ramesh and Mikhail Pavlov and Gabriel Goh and Scott Gray and Chelsea Voss and Alec Radford and Mark Chen and Ilya Sutskever},
year = {2021},
eprint = {2102.12092},
archivePrefix = {arXiv},
primaryClass = {cs.CV}
}
@misc{esser2021taming,
title = {Taming Transformers for High-Resolution Image Synthesis},
author = {Patrick Esser and Robin Rombach and BjΓΆrn Ommer},
year = {2021},
eprint = {2012.09841},
archivePrefix = {arXiv},
primaryClass = {cs.CV}
}
Python
100.0%