tyshiwo1/Awesome-Visual-Tokenizer

Awesome Visual Tokenizers/Autoencoders

20

14 commits

updated Nov 19, 2025

See the code

README

Awesome-Visual-Tokenizer

Categorize by Tokenizer Architecture:

Image Tokenizers

Continuous 2D CNN Image Tokenizer

  • Diagnosing and enhancing VAE models

  • Deep compression autoencoder for efficient high-resolution diffusion models

  • DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space

Discrete 2D CNN Image Tokenizer

  • Taming transformers for high-resolution image synthesis

    • Esser, P., Rombach, R., Ommer, B. et al.
    • CVPR 2021
    • 2021
  • Autoregressive image generation using residual quantization

    • Lee, D., Kim, C., Kim, S. et al.
    • CVPR 2022
    • 2022
  • Finite scalar quantization: Vq-vae made simple

  • Language Model Beats Diffusion--Tokenizer is Key to Visual Generation

  • Magvit: Masked generative video transformer

    • Yu, L., Cheng, Y., Sohn, K. et al.
    • CVPR 2023
    • 2023
  • Emu3: Next-Token Prediction is All You Need

  • Visual autoregressive modeling: Scalable image generation via next-scale prediction

  • Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

  • Open-magvit2: An open-source project toward democratizing auto-regressive visual generation

  • Quantize-then-Rectify: Efficient VQ-VAE Training

  • Image and video tokenization with binary spherical quantization

  • End-to-end vision tokenizer tuning

2D Transformer Image Tokenizer

  • Vector-quantized Image Modeling with Improved VQGAN

  • Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

  • Learnings from scaling visual tokenizers for reconstruction and generation

  • MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

  • Unitok: A unified tokenizer for visual generation and understanding

  • UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

  • Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation

    • Xiong, T., Liew, J.H., Huang, Z. et al.
    • ICCV 2025
    • 2025
  • Latent Diffusion Model without Variational Autoencoder

  • Diffusion Transformers with Representation Autoencoders

  • Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models

1D Image Tokenizer

  • An image is worth 32 tokens for reconstruction and generation

  • Elastictok: Adaptive tokenization for image and video

  • Latent Denoising Makes Good Visual Tokenizers

  • Language-guided image tokenization for generation

    • Zha, K., Yu, L., Fathi, A. et al.
    • CVPR 2025
    • 2025
  • Highly Compressed Tokenizer Can Generate Without Training

  • Masked autoencoders are effective tokenizers for diffusion models

    • Chen, H., Han, Y., Chen, F. et al.
    • ICML 2025
    • 2025
  • Softvq-vae: Efficient 1-dimensional continuous tokenizer

    • Chen, H., Wang, Z., Li, X. et al.
    • CVPR 2025
    • 2025
  • Adaptive length image tokenization via recurrent allocation

  • Democratizing text-to-image masked generative models with compact text-aware one-dimensional tokens

  • FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

    • Bachmann, R., Allardice, J., Mizrahi, D. et al.
    • ICML 2025
    • 2025
  • Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation

    • Xiong, T., Liew, J.H., Huang, Z. et al.
    • ICCV 2025
    • 2025
  • DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction

  • One-d-piece: Image tokenizer meets quality-controllable compression

  • Image Tokenizer Needs Post-Training

  • AliTok: Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model

Image Tokenizer with Generative Decoder

  • Diffusion Autoencoders: Toward a Meaningful and Decodable Representation

    • Preechakul, N., Chatthee, N., Wizadwongsa, S., Suwajanakorn, S.
    • CVPR 2022
    • 2022
  • epsilon-vae: Denoising as visual decoding

  • Diffusion autoencoders are scalable image tokenizers

  • FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

    • Bachmann, R., Allardice, J., Mizrahi, D. et al.
    • ICML 2025
    • 2025
  • D-AR: Diffusion via Autoregressive Models

  • Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

    • Pan, K., Lin, W., Yue, Z. et al.
    • CVPR 2025
    • 2025
  • Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization

  • " Principal Components" Enable A New Language of Images

Video Tokenizers

Continuous 3D CNN Video Tokenizer

  • Video generation models as world simulators

    • Brooks, T., Peebles, B., Holmes, C. et al.
    • The SORA
    • 2024
  • Cogvideox: Text-to-video diffusion models with an expert transformer

  • Open-sora plan: Open-source large video generation model

  • Hunyuanvideo: A systematic framework for large video generative models

  • Wan: Open and advanced large-scale video generative models

  • Ltx-video: Realtime video latent diffusion

  • Cosmos world foundation model platform for physical ai

  • LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models

  • Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model

    • Li, Z., Lin, B., Ye, Y. et al.
    • CVPR 2025
    • 2025
  • Step-video-t2v technical report: The practice, challenges, and future of video foundation model

  • Waver: Wave Your Way to Lifelike Video Generation

  • Vidtwin: Video vae with decoupled structure and dynamics

    • Wang, Y., Guo, J., Xie, X. et al.
    • CVPR 2025
    • 2025

Discrete 3D CNN Video Tokenizer

  • Language Model Beats Diffusion--Tokenizer is Key to Visual Generation

  • Videopoet: A large language model for zero-shot video generation

  • Loong: Generating minute-level long videos with autoregressive language models

  • Vidtok: A versatile and open-source video tokenizer

  • Omnitokenizer: A joint image-video tokenizer for visual generation

  • Cosmos world foundation model platform for physical ai

3D Transformer Video Tokenizer

  • MAGI-1: Autoregressive Video Generation at Scale

  • AToken: A Unified Tokenizer for Vision

  • Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

1D Video Tokenizer

  • Larp: Tokenizing videos with a learned autoregressive generative prior

  • Learning 1D causal visual representation with de-focus attention networks

Video Tokenizer with Generative Decoder

  • Rethinking video tokenization: A conditioned diffusion-based approach

  • REGEN: Learning Compact Video Embedding with (Re-) Generative Decoder

  • Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

Others

  • Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
    • Wang, Y., Lin, Z., Teng, Y. et al.
    • ICCV2025
    • 2025

Contributors

tyshiwo1

14 commits

tyshiwo1/Awesome-Visual-Tokenizer

Awesome Visual Tokenizers/Autoencoders

20

14 commits

updated Nov 19, 2025

See the code

README

Awesome-Visual-Tokenizer

Categorize by Tokenizer Architecture:

Image Tokenizers

Continuous 2D CNN Image Tokenizer

  • Diagnosing and enhancing VAE models

  • Deep compression autoencoder for efficient high-resolution diffusion models

  • DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space

Discrete 2D CNN Image Tokenizer

  • Taming transformers for high-resolution image synthesis

    • Esser, P., Rombach, R., Ommer, B. et al.
    • CVPR 2021
    • 2021
  • Autoregressive image generation using residual quantization

    • Lee, D., Kim, C., Kim, S. et al.
    • CVPR 2022
    • 2022
  • Finite scalar quantization: Vq-vae made simple

  • Language Model Beats Diffusion--Tokenizer is Key to Visual Generation

  • Magvit: Masked generative video transformer

    • Yu, L., Cheng, Y., Sohn, K. et al.
    • CVPR 2023
    • 2023
  • Emu3: Next-Token Prediction is All You Need

  • Visual autoregressive modeling: Scalable image generation via next-scale prediction

  • Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

  • Open-magvit2: An open-source project toward democratizing auto-regressive visual generation

  • Quantize-then-Rectify: Efficient VQ-VAE Training

  • Image and video tokenization with binary spherical quantization

  • End-to-end vision tokenizer tuning

2D Transformer Image Tokenizer

  • Vector-quantized Image Modeling with Improved VQGAN

  • Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

  • Learnings from scaling visual tokenizers for reconstruction and generation

  • MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

  • Unitok: A unified tokenizer for visual generation and understanding

  • UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

  • Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation

    • Xiong, T., Liew, J.H., Huang, Z. et al.
    • ICCV 2025
    • 2025
  • Latent Diffusion Model without Variational Autoencoder

  • Diffusion Transformers with Representation Autoencoders

  • Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models

1D Image Tokenizer

  • An image is worth 32 tokens for reconstruction and generation

  • Elastictok: Adaptive tokenization for image and video

  • Latent Denoising Makes Good Visual Tokenizers

  • Language-guided image tokenization for generation

    • Zha, K., Yu, L., Fathi, A. et al.
    • CVPR 2025
    • 2025
  • Highly Compressed Tokenizer Can Generate Without Training

  • Masked autoencoders are effective tokenizers for diffusion models

    • Chen, H., Han, Y., Chen, F. et al.
    • ICML 2025
    • 2025
  • Softvq-vae: Efficient 1-dimensional continuous tokenizer

    • Chen, H., Wang, Z., Li, X. et al.
    • CVPR 2025
    • 2025
  • Adaptive length image tokenization via recurrent allocation

  • Democratizing text-to-image masked generative models with compact text-aware one-dimensional tokens

  • FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

    • Bachmann, R., Allardice, J., Mizrahi, D. et al.
    • ICML 2025
    • 2025
  • Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation

    • Xiong, T., Liew, J.H., Huang, Z. et al.
    • ICCV 2025
    • 2025
  • DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction

  • One-d-piece: Image tokenizer meets quality-controllable compression

  • Image Tokenizer Needs Post-Training

  • AliTok: Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model

Image Tokenizer with Generative Decoder

  • Diffusion Autoencoders: Toward a Meaningful and Decodable Representation

    • Preechakul, N., Chatthee, N., Wizadwongsa, S., Suwajanakorn, S.
    • CVPR 2022
    • 2022
  • epsilon-vae: Denoising as visual decoding

  • Diffusion autoencoders are scalable image tokenizers

  • FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

    • Bachmann, R., Allardice, J., Mizrahi, D. et al.
    • ICML 2025
    • 2025
  • D-AR: Diffusion via Autoregressive Models

  • Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

    • Pan, K., Lin, W., Yue, Z. et al.
    • CVPR 2025
    • 2025
  • Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization

  • " Principal Components" Enable A New Language of Images

Video Tokenizers

Continuous 3D CNN Video Tokenizer

  • Video generation models as world simulators

    • Brooks, T., Peebles, B., Holmes, C. et al.
    • The SORA
    • 2024
  • Cogvideox: Text-to-video diffusion models with an expert transformer

  • Open-sora plan: Open-source large video generation model

  • Hunyuanvideo: A systematic framework for large video generative models

  • Wan: Open and advanced large-scale video generative models

  • Ltx-video: Realtime video latent diffusion

  • Cosmos world foundation model platform for physical ai

  • LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models

  • Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model

    • Li, Z., Lin, B., Ye, Y. et al.
    • CVPR 2025
    • 2025
  • Step-video-t2v technical report: The practice, challenges, and future of video foundation model

  • Waver: Wave Your Way to Lifelike Video Generation

  • Vidtwin: Video vae with decoupled structure and dynamics

    • Wang, Y., Guo, J., Xie, X. et al.
    • CVPR 2025
    • 2025

Discrete 3D CNN Video Tokenizer

  • Language Model Beats Diffusion--Tokenizer is Key to Visual Generation

  • Videopoet: A large language model for zero-shot video generation

  • Loong: Generating minute-level long videos with autoregressive language models

  • Vidtok: A versatile and open-source video tokenizer

  • Omnitokenizer: A joint image-video tokenizer for visual generation

  • Cosmos world foundation model platform for physical ai

3D Transformer Video Tokenizer

  • MAGI-1: Autoregressive Video Generation at Scale

  • AToken: A Unified Tokenizer for Vision

  • Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

1D Video Tokenizer

  • Larp: Tokenizing videos with a learned autoregressive generative prior

  • Learning 1D causal visual representation with de-focus attention networks

Video Tokenizer with Generative Decoder

  • Rethinking video tokenization: A conditioned diffusion-based approach

  • REGEN: Learning Compact Video Embedding with (Re-) Generative Decoder

  • Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion

Others

  • Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
    • Wang, Y., Lin, Z., Teng, Y. et al.
    • ICCV2025
    • 2025

Contributors

tyshiwo1

14 commits