This repo aims to provide a comprehensive comparison on open-sourced discrete visual tokenizers under a fair setting: same resoution, same data, same evaluation code.
We use Python 3.10 and Pytorch 2.1.2, with CUDA Version 12.2. Please install the required packages using the following command:
pip install -e .
We use a subset of laion-aesthetics-12m-umap and OpenVid-1M as the test set to benchmark all the available tokenizers. You can download the data directly from here.
The comparison of image and video reconstruction performance is shown below:
| Name | Compression | PSNR | SSIM | rFID | FPS2 |
|---|---|---|---|---|---|
| OmniTokenizer1 | 8x8 | 32.16 | 0.86 | 2.02 | 121.24 |
| Cosmos-DI | 8x8 | 31.04 | 0.83 | 0.95 | 612.91 |
| Emu3 | 8x8 | 31.51 | 0.83 | 0.57 | 12.72 |
| LlamaGen | 16x16 | 26.98 | 0.71 | 1.64 | 198.06 |
| Show-O | 16x16 | 27.03 | 0.71 | 1.54 | 129.62 |
| Tiktok | 1D | 27.48 | 0.72 | 1.30 | 303.06 |
| Name | Compression | PSNR | SSIM | rFID | FPS |
|---|---|---|---|---|---|
| OmniTokenizer | 2x8x8 | 33.51 | 0.93 | 4.10 | 67.08 |
| Cosmos-DV | 4x8x8 | 34.54 | 0.94 | 4.19 | 76.46 |
| Emu3 | 4x8x8 | 33.44 | 0.92 | 8.15 | 3.24 |
1: OmniTokenizer is the only tokenizer designed for joint image and video tokenization. Here we release an any-res checkpoint of OmniTokenizer, meaning we train the model with 2D Rope on images / videos with different aspect ratios and resolutions.
2: we denote the throughput of tokenizers as FPS, images (videos) per second.
Please run the following code for reconstruction:
torchrun \
--nnodes=1 --nproc_per_node=4 --master_port 23456 \
eval/reconstruct.py \
--vq_model_type "omnitokenizer" \
--vq_model_ckpt "./checkpoints/omnitokenizer_rq_code16384_down16_joint2.ckpt" \
--dataset_type "video" \
--dataset_name "openvid" \
--save_dir "./tokenizers" \
--video_path "path_to_dir/tokenizer_bench/openvid.json" \
--video_folder "" \
--resolution 256 \
--video_fps 16 \
--sequence_length 17
Please specify the tokenizer type with --vq_model_type, the ckpt with --vq_model_ckpt. The above command will automatically save the reconstructed results under --save_dir. We provide the script to inference different tokenizers under ./scripts/eval.
After this, you may run the following command to obtain the reconstruction metrics:
# For image reconstruction:
python eval/calculate_image.py /path_to_gt_dir /path_to_recon_dir
# For video reconstruction:
python eval/calculate_video.py /path_to_video_dir
We also provide a script to facilitate the token extraction using different tokenizers in an offline manner in ./eval/extract_token.py.
We show the 256x256 image reconstruction results below:
From left to right, we compare OmniTokenizer, Cosmos-Tokenizer, and Emu3:




If you consider our work useful, please consider citing our paper using:
@inproceedings{wang2024omnitokenizer,
title={OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation},
author={Wang, Junke and Jiang, Yi and Yuan, Zehuan and Peng, Binyue and Wu, Zuxuan and Jiang, Yu-Gang},
booktitle={NeurIPS},
year={2024}
}
Thanks OmniTokenizer, LlamaGen, Cosmos-Tokenizer, Emu3, Show-O, and TikTok for their great work. We also borrow several functions from Reducio for evaluation.
Python
99.4%
This repo aims to provide a comprehensive comparison on open-sourced discrete visual tokenizers under a fair setting: same resoution, same data, same evaluation code.
We use Python 3.10 and Pytorch 2.1.2, with CUDA Version 12.2. Please install the required packages using the following command:
pip install -e .
We use a subset of laion-aesthetics-12m-umap and OpenVid-1M as the test set to benchmark all the available tokenizers. You can download the data directly from here.
The comparison of image and video reconstruction performance is shown below:
| Name | Compression | PSNR | SSIM | rFID | FPS2 |
|---|---|---|---|---|---|
| OmniTokenizer1 | 8x8 | 32.16 | 0.86 | 2.02 | 121.24 |
| Cosmos-DI | 8x8 | 31.04 | 0.83 | 0.95 | 612.91 |
| Emu3 | 8x8 | 31.51 | 0.83 | 0.57 | 12.72 |
| LlamaGen | 16x16 | 26.98 | 0.71 | 1.64 | 198.06 |
| Show-O | 16x16 | 27.03 | 0.71 | 1.54 | 129.62 |
| Tiktok | 1D | 27.48 | 0.72 | 1.30 | 303.06 |
| Name | Compression | PSNR | SSIM | rFID | FPS |
|---|---|---|---|---|---|
| OmniTokenizer | 2x8x8 | 33.51 | 0.93 | 4.10 | 67.08 |
| Cosmos-DV | 4x8x8 | 34.54 | 0.94 | 4.19 | 76.46 |
| Emu3 | 4x8x8 | 33.44 | 0.92 | 8.15 | 3.24 |
1: OmniTokenizer is the only tokenizer designed for joint image and video tokenization. Here we release an any-res checkpoint of OmniTokenizer, meaning we train the model with 2D Rope on images / videos with different aspect ratios and resolutions.
2: we denote the throughput of tokenizers as FPS, images (videos) per second.
Please run the following code for reconstruction:
torchrun \
--nnodes=1 --nproc_per_node=4 --master_port 23456 \
eval/reconstruct.py \
--vq_model_type "omnitokenizer" \
--vq_model_ckpt "./checkpoints/omnitokenizer_rq_code16384_down16_joint2.ckpt" \
--dataset_type "video" \
--dataset_name "openvid" \
--save_dir "./tokenizers" \
--video_path "path_to_dir/tokenizer_bench/openvid.json" \
--video_folder "" \
--resolution 256 \
--video_fps 16 \
--sequence_length 17
Please specify the tokenizer type with --vq_model_type, the ckpt with --vq_model_ckpt. The above command will automatically save the reconstructed results under --save_dir. We provide the script to inference different tokenizers under ./scripts/eval.
After this, you may run the following command to obtain the reconstruction metrics:
# For image reconstruction:
python eval/calculate_image.py /path_to_gt_dir /path_to_recon_dir
# For video reconstruction:
python eval/calculate_video.py /path_to_video_dir
We also provide a script to facilitate the token extraction using different tokenizers in an offline manner in ./eval/extract_token.py.
We show the 256x256 image reconstruction results below:
From left to right, we compare OmniTokenizer, Cosmos-Tokenizer, and Emu3:




If you consider our work useful, please consider citing our paper using:
@inproceedings{wang2024omnitokenizer,
title={OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation},
author={Wang, Junke and Jiang, Yi and Yuan, Zehuan and Peng, Binyue and Wu, Zuxuan and Jiang, Yu-Gang},
booktitle={NeurIPS},
year={2024}
}
Thanks OmniTokenizer, LlamaGen, Cosmos-Tokenizer, Emu3, Show-O, and TikTok for their great work. We also borrow several functions from Reducio for evaluation.
Python
99.4%