Categorize by Tokenizer Architecture:
Diagnosing and enhancing VAE models
Deep compression autoencoder for efficient high-resolution diffusion models
DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
Taming transformers for high-resolution image synthesis
Autoregressive image generation using residual quantization
Finite scalar quantization: Vq-vae made simple
Language Model Beats Diffusion--Tokenizer is Key to Visual Generation
Magvit: Masked generative video transformer
Emu3: Next-Token Prediction is All You Need
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation
Quantize-then-Rectify: Efficient VQ-VAE Training
Image and video tokenization with binary spherical quantization
End-to-end vision tokenizer tuning
Vector-quantized Image Modeling with Improved VQGAN
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
Learnings from scaling visual tokenizers for reconstruction and generation
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
Unitok: A unified tokenizer for visual generation and understanding
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation
Latent Diffusion Model without Variational Autoencoder
Diffusion Transformers with Representation Autoencoders
Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
An image is worth 32 tokens for reconstruction and generation
Elastictok: Adaptive tokenization for image and video
Latent Denoising Makes Good Visual Tokenizers
Language-guided image tokenization for generation
Highly Compressed Tokenizer Can Generate Without Training
Masked autoencoders are effective tokenizers for diffusion models
Softvq-vae: Efficient 1-dimensional continuous tokenizer
Adaptive length image tokenization via recurrent allocation
Democratizing text-to-image masked generative models with compact text-aware one-dimensional tokens
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation
DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
One-d-piece: Image tokenizer meets quality-controllable compression
Image Tokenizer Needs Post-Training
AliTok: Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
Diffusion Autoencoders: Toward a Meaningful and Decodable Representation
epsilon-vae: Denoising as visual decoding
Diffusion autoencoders are scalable image tokenizers
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
D-AR: Diffusion via Autoregressive Models
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization
" Principal Components" Enable A New Language of Images
Video generation models as world simulators
Cogvideox: Text-to-video diffusion models with an expert transformer
Open-sora plan: Open-source large video generation model
Hunyuanvideo: A systematic framework for large video generative models
Wan: Open and advanced large-scale video generative models
Ltx-video: Realtime video latent diffusion
Cosmos world foundation model platform for physical ai
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model
Step-video-t2v technical report: The practice, challenges, and future of video foundation model
Waver: Wave Your Way to Lifelike Video Generation
Vidtwin: Video vae with decoupled structure and dynamics
Language Model Beats Diffusion--Tokenizer is Key to Visual Generation
Videopoet: A large language model for zero-shot video generation
Loong: Generating minute-level long videos with autoregressive language models
Vidtok: A versatile and open-source video tokenizer
Omnitokenizer: A joint image-video tokenizer for visual generation
Cosmos world foundation model platform for physical ai
MAGI-1: Autoregressive Video Generation at Scale
AToken: A Unified Tokenizer for Vision
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
Larp: Tokenizing videos with a learned autoregressive generative prior
Learning 1D causal visual representation with de-focus attention networks
Rethinking video tokenization: A conditioned diffusion-based approach
REGEN: Learning Compact Video Embedding with (Re-) Generative Decoder
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
14 commits
Categorize by Tokenizer Architecture:
Diagnosing and enhancing VAE models
Deep compression autoencoder for efficient high-resolution diffusion models
DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
Taming transformers for high-resolution image synthesis
Autoregressive image generation using residual quantization
Finite scalar quantization: Vq-vae made simple
Language Model Beats Diffusion--Tokenizer is Key to Visual Generation
Magvit: Masked generative video transformer
Emu3: Next-Token Prediction is All You Need
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation
Quantize-then-Rectify: Efficient VQ-VAE Training
Image and video tokenization with binary spherical quantization
End-to-end vision tokenizer tuning
Vector-quantized Image Modeling with Improved VQGAN
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
Learnings from scaling visual tokenizers for reconstruction and generation
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
Unitok: A unified tokenizer for visual generation and understanding
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation
Latent Diffusion Model without Variational Autoencoder
Diffusion Transformers with Representation Autoencoders
Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
An image is worth 32 tokens for reconstruction and generation
Elastictok: Adaptive tokenization for image and video
Latent Denoising Makes Good Visual Tokenizers
Language-guided image tokenization for generation
Highly Compressed Tokenizer Can Generate Without Training
Masked autoencoders are effective tokenizers for diffusion models
Softvq-vae: Efficient 1-dimensional continuous tokenizer
Adaptive length image tokenization via recurrent allocation
Democratizing text-to-image masked generative models with compact text-aware one-dimensional tokens
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Gigatok: Scaling visual tokenizers to 3 billion parameters for autoregressive image generation
DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
One-d-piece: Image tokenizer meets quality-controllable compression
Image Tokenizer Needs Post-Training
AliTok: Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
Diffusion Autoencoders: Toward a Meaningful and Decodable Representation
epsilon-vae: Denoising as visual decoding
Diffusion autoencoders are scalable image tokenizers
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
D-AR: Diffusion via Autoregressive Models
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization
" Principal Components" Enable A New Language of Images
Video generation models as world simulators
Cogvideox: Text-to-video diffusion models with an expert transformer
Open-sora plan: Open-source large video generation model
Hunyuanvideo: A systematic framework for large video generative models
Wan: Open and advanced large-scale video generative models
Ltx-video: Realtime video latent diffusion
Cosmos world foundation model platform for physical ai
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models
Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model
Step-video-t2v technical report: The practice, challenges, and future of video foundation model
Waver: Wave Your Way to Lifelike Video Generation
Vidtwin: Video vae with decoupled structure and dynamics
Language Model Beats Diffusion--Tokenizer is Key to Visual Generation
Videopoet: A large language model for zero-shot video generation
Loong: Generating minute-level long videos with autoregressive language models
Vidtok: A versatile and open-source video tokenizer
Omnitokenizer: A joint image-video tokenizer for visual generation
Cosmos world foundation model platform for physical ai
MAGI-1: Autoregressive Video Generation at Scale
AToken: A Unified Tokenizer for Vision
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
Larp: Tokenizing videos with a learned autoregressive generative prior
Learning 1D causal visual representation with de-focus attention networks
Rethinking video tokenization: A conditioned diffusion-based approach
REGEN: Learning Compact Video Embedding with (Re-) Generative Decoder
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
14 commits