wdrink/OpenTokenizer

Python

21

0 commits

updated Jan 17, 2025

See the code

README

OpenTokenizer: A Comprehensive Comparision on Open-sourced Visual Tokenizers

This repo aims to provide a comprehensive comparison on open-sourced discrete visual tokenizers under a fair setting: same resoution, same data, same evaluation code.

Setup

We use Python 3.10 and Pytorch 2.1.2, with CUDA Version 12.2. Please install the required packages using the following command:

pip install -e .

Benchmark

We use a subset of laion-aesthetics-12m-umap and OpenVid-1M as the test set to benchmark all the available tokenizers. You can download the data directly from here.

The comparison of image and video reconstruction performance is shown below:

NameCompressionPSNRSSIMrFIDFPS2
OmniTokenizer18x832.160.862.02121.24
Cosmos-DI8x831.040.830.95612.91
Emu38x831.510.830.5712.72
LlamaGen16x1626.980.711.64198.06
Show-O16x1627.030.711.54129.62
Tiktok1D27.480.721.30303.06
NameCompressionPSNRSSIMrFIDFPS
OmniTokenizer2x8x833.510.934.1067.08
Cosmos-DV4x8x834.540.944.1976.46
Emu34x8x833.440.928.153.24

1: OmniTokenizer is the only tokenizer designed for joint image and video tokenization. Here we release an any-res checkpoint of OmniTokenizer, meaning we train the model with 2D Rope on images / videos with different aspect ratios and resolutions.

2: we denote the throughput of tokenizers as FPS, images (videos) per second.

Usage

Please run the following code for reconstruction:

torchrun \
--nnodes=1 --nproc_per_node=4 --master_port 23456 \
eval/reconstruct.py \
    --vq_model_type "omnitokenizer" \
    --vq_model_ckpt "./checkpoints/omnitokenizer_rq_code16384_down16_joint2.ckpt" \
    --dataset_type "video" \
    --dataset_name "openvid" \
    --save_dir "./tokenizers" \
    --video_path "path_to_dir/tokenizer_bench/openvid.json" \
    --video_folder "" \
    --resolution 256 \
    --video_fps 16 \
    --sequence_length 17

Please specify the tokenizer type with --vq_model_type, the ckpt with --vq_model_ckpt. The above command will automatically save the reconstructed results under --save_dir. We provide the script to inference different tokenizers under ./scripts/eval.

After this, you may run the following command to obtain the reconstruction metrics:

# For image reconstruction:

python eval/calculate_image.py /path_to_gt_dir /path_to_recon_dir

# For video reconstruction:

python eval/calculate_video.py /path_to_video_dir

Offline Code Extraction

We also provide a script to facilitate the token extraction using different tokenizers in an offline manner in ./eval/extract_token.py.

Results

We show the 256x256 image reconstruction results below:

From left to right, we compare OmniTokenizer, Cosmos-Tokenizer, and Emu3:

To Do List

  • open source the evaluation code
  • support the calculation of different metrics
  • open source the training code

Citation

If you consider our work useful, please consider citing our paper using:

@inproceedings{wang2024omnitokenizer,
  title={OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation},
  author={Wang, Junke and Jiang, Yi and Yuan, Zehuan and Peng, Binyue and Wu, Zuxuan and Jiang, Yu-Gang},
  booktitle={NeurIPS},
  year={2024}
}

Acknowledgement

Thanks OmniTokenizer, LlamaGen, Cosmos-Tokenizer, Emu3, Show-O, and TikTok for their great work. We also borrow several functions from Reducio for evaluation.

wdrink/OpenTokenizer

Python

21

0 commits

updated Jan 17, 2025

See the code

README

OpenTokenizer: A Comprehensive Comparision on Open-sourced Visual Tokenizers

This repo aims to provide a comprehensive comparison on open-sourced discrete visual tokenizers under a fair setting: same resoution, same data, same evaluation code.

Setup

We use Python 3.10 and Pytorch 2.1.2, with CUDA Version 12.2. Please install the required packages using the following command:

pip install -e .

Benchmark

We use a subset of laion-aesthetics-12m-umap and OpenVid-1M as the test set to benchmark all the available tokenizers. You can download the data directly from here.

The comparison of image and video reconstruction performance is shown below:

NameCompressionPSNRSSIMrFIDFPS2
OmniTokenizer18x832.160.862.02121.24
Cosmos-DI8x831.040.830.95612.91
Emu38x831.510.830.5712.72
LlamaGen16x1626.980.711.64198.06
Show-O16x1627.030.711.54129.62
Tiktok1D27.480.721.30303.06
NameCompressionPSNRSSIMrFIDFPS
OmniTokenizer2x8x833.510.934.1067.08
Cosmos-DV4x8x834.540.944.1976.46
Emu34x8x833.440.928.153.24

1: OmniTokenizer is the only tokenizer designed for joint image and video tokenization. Here we release an any-res checkpoint of OmniTokenizer, meaning we train the model with 2D Rope on images / videos with different aspect ratios and resolutions.

2: we denote the throughput of tokenizers as FPS, images (videos) per second.

Usage

Please run the following code for reconstruction:

torchrun \
--nnodes=1 --nproc_per_node=4 --master_port 23456 \
eval/reconstruct.py \
    --vq_model_type "omnitokenizer" \
    --vq_model_ckpt "./checkpoints/omnitokenizer_rq_code16384_down16_joint2.ckpt" \
    --dataset_type "video" \
    --dataset_name "openvid" \
    --save_dir "./tokenizers" \
    --video_path "path_to_dir/tokenizer_bench/openvid.json" \
    --video_folder "" \
    --resolution 256 \
    --video_fps 16 \
    --sequence_length 17

Please specify the tokenizer type with --vq_model_type, the ckpt with --vq_model_ckpt. The above command will automatically save the reconstructed results under --save_dir. We provide the script to inference different tokenizers under ./scripts/eval.

After this, you may run the following command to obtain the reconstruction metrics:

# For image reconstruction:

python eval/calculate_image.py /path_to_gt_dir /path_to_recon_dir

# For video reconstruction:

python eval/calculate_video.py /path_to_video_dir

Offline Code Extraction

We also provide a script to facilitate the token extraction using different tokenizers in an offline manner in ./eval/extract_token.py.

Results

We show the 256x256 image reconstruction results below:

From left to right, we compare OmniTokenizer, Cosmos-Tokenizer, and Emu3:

To Do List

  • open source the evaluation code
  • support the calculation of different metrics
  • open source the training code

Citation

If you consider our work useful, please consider citing our paper using:

@inproceedings{wang2024omnitokenizer,
  title={OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation},
  author={Wang, Junke and Jiang, Yi and Yuan, Zehuan and Peng, Binyue and Wu, Zuxuan and Jiang, Yu-Gang},
  booktitle={NeurIPS},
  year={2024}
}

Acknowledgement

Thanks OmniTokenizer, LlamaGen, Cosmos-Tokenizer, Emu3, Show-O, and TikTok for their great work. We also borrow several functions from Reducio for evaluation.

Languages

Python

99.4%