ATH-MaaS/Awesome-Unified-Multimodal-Models

Awesome Unified Multimodal Models

1,321

32 commits

updated Mar 24, 2026

See the code

README

Awesome Unified Multimodal Models

📚Survey • 🤗 HF Repo

Figure 1: Timeline of Publicly Available and Unavailable Unified Multimodal Models. The models are categorized by their release years, from 2023 to 2025. Models underlined in the diagram represent any-to-any multimodal models, capable of handling inputs or outputs beyond text and image, such as audio, video, and speech. The timeline highlights the rapid growth in this field.

🔥 We are hiring!

We are looking for both interns and full-time researchers to join our team, focusing on multimodal understanding, generation, reasoning, AI agents, and unified multimodal models. If you are interested in exploring these exciting areas, please reach out to us at qingguo.cqg@alibaba-inc.com.

👉 What is This Repo for?

This repository provides a comprehensive collection of resources related to unified multimodal models, featuring:

  • A survey of advances, challenges, and timelines for unified models
  • Categorized lists of diffusion-based, autoregressive (MLLM), and hybrid architectures for unified image–text understanding and generation
  • Benchmarks for evaluating multimodal comprehension, image generation, and interleaved image–text tasks
  • Representative datasets covering multimodal understanding, text-to-image synthesis, image editing, and interleaved interactions

Designed to help researchers and practitioners explore, compare, and build state-of-the-art unified multimodal systems.

Awesome Papers & Datasets

Text-and-Image Unified Models

Figure 2: Classification of Unified Multimodal Understanding and Generation Models. The models are divided into three main categories based on their backbone architecture: Diffusion, MLLM (AR), and MLLM (AR + Diffusion). Each category is further subdivided according to the encoding strategy employed, including Pixel Encoding, Semantic Encoding, Learnable Query Encoding, and Hybrid Encoding. We illustrate the architectural variations within these categories and their corresponding encoder-decoder configurations.

Diffusion

MLLM AR

b-1: Pixel Encoding
NameTitleVenueDateCodeDemo
Emu3.5Emu3.5: Native Multimodal Models are World Learners GitHub Repo starsarXiv2025/10/30Github-
Uni-XUni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models GitHub Repo starsICLR2025/09/29Github-
OneCatOneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation GitHub Repo starsarXiv2025/09/03Github-
SelftokSelftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning GitHub Repo starsarXiv2025/05/12Github-
TokLIPTokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation GitHub Repo starsarXiv2025/05/08Github-
HarmonHarmonizing Visual Representations for Unified Multimodal Understanding and Generation GitHub Repo starsarXiv2025/03/27GithubDemo
UGenUGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary LearningarXiv2025/03/27--
SynerGen-VLSynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingarXiv2024/12/12--
LiquidLiquid: Language Models are Scalable and Unified Multi-modal Generators GitHub Repo starsarXiv2024/12/05GithubDemo
OrthusOrthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads GitHub Repo starsarXiv2024/11/28Github-
MMARMMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic ModelingarXiv2024/10/14--
Emu3Emu3: Next-Token Prediction is All You Need GitHub Repo starsarXiv2024/09/27GithubDemo
ANOLEANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation GitHub Repo starsarXiv2024/07/08Github-
ChameleonChameleon: Mixed-Modal Early-Fusion Foundation Models GitHub Repo starsarXiv2024/05/16Github-
LWMWorld Model on Million-Length Video And Language With Blockwise RingAttention GitHub Repo starsICLR2024/02/13Github-
b-2: Semantic Encoding
NameTitleVenueDateCodeDemo
MammothModa2MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and GenerationGitHub Repo starsarXiv2025/11/23GitHub-
Ming-UniVisionMing-UniVision: Joint Image Understanding and Generation with a Unified Continuous TokenizerGitHub Repo starsarXiv2025/10/08GitHub-
Bifrost-1Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP LatentsGitHub Repo starsarXiv2025/08/08GitHub-
Qwen-ImageQwen-Image Technical ReportGitHub Repo starsarXiv2025/08/04GitHubDemo
X-OmniX-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great AgainGitHub Repo starsarXiv2025/07/29GitHubDemo
Ovis-U1Ovis-U1 Technical ReportGitHub Repo starsarXiv2025/06/28GitHubDemo
UniCode$^2$UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and GenerationarXiv2025/06/20--
OmniGen2OmniGen2: Exploration to Advanced Multimodal Generation GitHub Repo starsarXiv2025/06/18GithubDemo
TarVision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GitHub Repo starsarXiv2025/06/18GithubDemo
UniForkUniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation GitHub Repo starsarXiv2025/06/17Github-
UniWorldUniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation GitHub Repo starsarXiv2025/06/03Github-
PiscesPisces: An Auto-regressive Foundation Model for Image Understanding and Generation GitHub Repo starsarXiv2025/06/10Github-
DualTokenDualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies GitHub Repo starsarXiv2025/03/18Github-
UniTokUniTok: A Unified Tokenizer for Visual Generation and Understanding GitHub Repo starsarXiv2025/02/27GithubDemo
QLIPQLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation GitHub Repo starsarXiv2025/02/05Github-
MetaMorphMetaMorph: Multimodal Understanding and Generation via Instruction Tuning GitHub Repo starsarXiv2024/12/18Github-
ILLUMEILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhancearXiv2024/12/09--
PUMAPUMA: Empowering Unified MLLM with Multi-granular Visual Generation GitHub Repo starsarXiv2024/10/17Github-
VILA-UVILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation GitHub Repo starsICLR2024/09/06GithubDemo
Mini-GeminiMini-Gemini: Mining the Potential of Multi-modality Vision Language Models GitHub Repo starsarXiv2024/03/27GithubDemo
MM-InterleavedMM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer GitHub Repo starsarXiv2024/01/18Github-
VL-GPTVL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation GitHub Repo starsarXiv2023/12/14Github-
Emu2Generative Multimodal Models are In-Context Learners GitHub Repo starsCVPR2023/12/10GithubDemo
DreamLLMDreamLLM: Synergistic Multimodal Comprehension and Creation GitHub Repo starsICLR2023/09/20Github-
LaVITUnified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization GitHub Repo starsICLR2023/09/09Github-
EmuEmu: Generative Pretraining in Multimodality GitHub Repo starsICLR2023/07/11GithubDemo
b-3: Learnable Query Encoding
b-4: Hybrid Encoding (Pseduo)
b-5: Hybrid Encoding (Joint)

MLLM AR-Diffusion

c-1: Pxiel Encoding
c-2: Hybrid Encoding (Pseduo)

Any-to-Any Multimodal models

Benchmark for Evaluation

Benchmarks on Understanding Tasks

Benchmarks on Image Generation Tasks

NamePaperVenueDateCode
GenExamGenExam: A Multidisciplinary Text-to-Image Exam Stararxiv2025/09/18Github
CVTGTextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes Stararxiv2025/08/05Github
OneIG-BenchOneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation Stararxiv2025/06/26Github
ComplexBench-EditComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies Stararxiv2025/06/15Github
EditInspectorEditInspector: A Benchmark for Evaluation of Text-Guided Image Editsarxiv2025/06/11-
ByteMorph-BenchByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions Stararxiv2025/06/03Github
RefEdit-BenchRefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions Stararxiv2025/06/03Github
WISEWISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Stararxiv2025/05/27Github
ImgEdit-BenchImgEdit: A Unified Image Editing Dataset and Benchmark Stararxiv2025/05/26Github
MMIG-BenchMMGen-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Stararxiv2025/05/26Github
KRIS-BenchKRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models StarNeurIPS 20252025/05/22Github
CompBenchCompBench: Benchmarking Complex Instruction-guided Image Editingarxiv2025/05/18-
WorldGenBenchWorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generationarxiv2025/05/02HuggingFace
GEdit-BenchStep1X-Edit: A Practical Framework for General Image Editing StararXiv2025/04/28Github
DreamBench++DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation StarICLR2025/03/09Github
T2I-CompBench++T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation StarTPAMI2025/03/08Github
IE-BenchIE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception AlignmentarXiv2025/01/17-
AnyEditAnyEdit: Mastering Unified High-Quality Image Editing for Any Idea StarCVPR2024/11/24Github
I2EBenchI2EBench: A Comprehensive Benchmark for Instruction-based Image Editing StarNeurIPS2024/08/26Github
ConceptMixConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty StarNeurIPS2024/08/26Github
GenAI-BenchGenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation StarCVPR2024/06/19Github
Commonsense-T2ICommonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? StarCOLM2024/06/11Github
HQ-EditHQ-Edit: A High-Quality Dataset for Instruction-based Image Editing StarICLR2024/04/15Github
VQAScoreEvaluating Text-to-Visual Generation with Image-to-Text Generation StarECCV2024/04/01Github
FlashEvalFlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models StarCVPR2024/03/25Github
DPG-BenchELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment Stararxiv2024/03/08Github
Reason-EditSmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models StarCVPR2023/12/11Github
Emu EditEmu Edit: Precise Image Editing via Recognition and Generation TasksCVPR2023/11/16HuggingFace
HEIMHolistic Evaluation of Text-To-Image Models StarNeurIPS2023/11/07Github
DSG-1kDavidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation StarICLR2023/10/27Github
GenEvalGenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment StarNeurIPS2023/10/17Github
EditValEditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods StararXiv2023/10/03Github
T2I-CompBenchT2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation StarNeurIPS2023/07/12Github
DreamSimDreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data StarNeurIPS2023/06/15Github
MagicBrushMagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing StarNeurIPS2023/06/16Github
MultiGen-20MUniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild StarNeurIPS2023/05/18Github
HRS-BenchHRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models StarICCV2023/04/11Github
TIFATIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering StarICCV2023/03/21Github
EditBenchImagen Editor and EditBench: Advancing and Evaluating Text-Guided Image InpaintingCVPR2022/12/13ProjectPage
PartiPromptsScaling Autoregressive Models for Content-Rich Text-to-Image Generation StarTMLR2022/06/22Github
DrawBenchPhotorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingNeurIPS2022/05/23ProjectPage
PaintSkillsDALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models StarICCV2022/02/08Github

Benchmarks on Interleaved / Compositional / Other Tasks

Dataset

Multimodal Understanding

Text-to-Image

DatasetSamplesPaperVenueDate
FLUX-Reason-6M6MFLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive BenchmarkarXiv2025/09/11
Echo-4o-Image106KEcho-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image GenerationarXiv2025/08/13
Poster100K100KPosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified FrameworkarXiv2025/06/12
Text-Render-2M2MPosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified FrameworkarXiv2025/06/12
ShareGPT-4o-Image45KShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image GenerationarXiv2025/06/22
BLIP3o-60k60KBLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and DatasetarXiv2025/05/14
TextAtlas5M5MTextAtlas5M: A Large-scale Dataset for Dense Text Image GenerationarXiv2025/02/11
EliGen TrainSet500KEliGen: Entity-Level Controlled Image Generation with Regional AttentionarXiv2025/01/02
PD12M12MPublic domain 12m: A highly aesthetic image-text dataset with novel governance mechanismsarXiv2024/10/30
SFHQ-T2I122K--2024/10/06
text-to-image-2M2M--2024/09/13
DenseFusion1MDensefusion-1m: Merging vision experts for comprehensive multimodal perceptionNeurIPS2024/07/11
Megalith10M--2024/07/01
PixelProse16MFrom pixels to prose: A large dataset of dense image captionsarXiv2024/06/14
DOCCI15KDOCCI: Descriptions of Connected and Contrasting ImagesECCV2024/04/30
CosmicMan-HQ 1.06MCosmicman: A text-to-image foundation model for humansCVPR2024/04/01
AnyWord-3M3MAnytext: Multilingual visual text generation and editingICLR2023/11/06
JourneyDB4MJourneyDB: A Benchmark for Generative Image UnderstandingNeurIPS2023/07/03
RenderedText12M--2023/06/30
Mario-10M10MTextdiffuser: Diffusion models as text paintersNeurIPS2023/05/18
SAM11MSegment AnythingICCV2023/04/05
LAION-Aesthetics120MLaion-5b: An open large-scale dataset for training next generation image-text modelsNeurIPS2022/08/16
CC-12M12MConceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual conceptsCVPR2021/02/17

Image Editing

DatasetSamplesPaperVenueDate
Pico-Banana-400K400KPico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingarXiv2025/10/22
X2Edit3.7MX2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation LearningarXiv2025/08/11
ShareGPT-4o-Image (Editing)46KShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image GenerationarXiv2025/06/22
ByteMorph-6M6MByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid MotionsarXiv2025/06/03
ImgEdit1.2MImgEdit: A Unified Image Editing Dataset and BenchmarkarXiv2025/05/26
RefEdit18KRefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model for Referring ExpressionarXiv2025/04/03
AnyEdit2.5MAnyedit: Mastering unified high-quality image editing for any ideaCVPR2024/11/24
OmniEdit1.2MOmniedit: Building image editing generalist models through specialist supervisionICLR2024/11/11
PromptFix1MPromptFix: You Prompt and We Fix the PhotoNeurIPS2024/09/19
UltraEdit4MUltraedit: Instruction-based fine-grained image editing at scaleNeurIPS2024/07/07
EditWorld8.6KEditWorld: Simulating World Dynamics for Instruction-Following Image EditingarXiv2024/06/23
SEED-Data-Edit3.7MSeed-data-edit technical report: A hybrid dataset for instructional image editingarXiv2024/05/07
HQ-Edit197KHq-edit: A high-quality dataset for instruction-based image editingarXiv2024/04/15
HIVE1.1MHIVE: Harnessing Human Feedback for Instructional Visual EditingarXiv2023/07/08
Magicbrush10KMagicbrush: A manually annotated dataset for instruction-guided image editingNeurIPS2023/06/16
InstructP2P313KInstructpix2pix: Learning to follow image editing instructionsCVPR2022/11/17

Interleaved Image-Text

Other Text-Image-to-Image

Applications and Opportunities

Citation

If you find this repo is helpful for your research, please cite our paper:

@article{zhang2025unified,
  title={Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities},
  author={Zhang, Xinjie and Guo, Jintao and Zhao, Shanshan and Fu, Minghao and Duan, Lunhao and Hu, Jiakui and Chng, Yong Xien and Wang, Guo-Hua and Chen, Qing-Guo and Xu, Zhao and Luo, Weihua and Zhang, Kaifu},
  journal={arXiv preprint arXiv:2505.02567},
  year={2025}
}
multimodal-large-language-models
multimodal-models
text-to-image-generation
unified-multimodal-models
vision-language-model

Contributors

suikong-zss

16 commits

sshan-zhao

10 commits

SxJyJay

2 commits

rollingsu

1 commits

ATH-MaaS/Awesome-Unified-Multimodal-Models

Awesome Unified Multimodal Models

1,321

32 commits

updated Mar 24, 2026

See the code

README

Awesome Unified Multimodal Models

📚Survey • 🤗 HF Repo

Figure 1: Timeline of Publicly Available and Unavailable Unified Multimodal Models. The models are categorized by their release years, from 2023 to 2025. Models underlined in the diagram represent any-to-any multimodal models, capable of handling inputs or outputs beyond text and image, such as audio, video, and speech. The timeline highlights the rapid growth in this field.

🔥 We are hiring!

We are looking for both interns and full-time researchers to join our team, focusing on multimodal understanding, generation, reasoning, AI agents, and unified multimodal models. If you are interested in exploring these exciting areas, please reach out to us at qingguo.cqg@alibaba-inc.com.

👉 What is This Repo for?

This repository provides a comprehensive collection of resources related to unified multimodal models, featuring:

  • A survey of advances, challenges, and timelines for unified models
  • Categorized lists of diffusion-based, autoregressive (MLLM), and hybrid architectures for unified image–text understanding and generation
  • Benchmarks for evaluating multimodal comprehension, image generation, and interleaved image–text tasks
  • Representative datasets covering multimodal understanding, text-to-image synthesis, image editing, and interleaved interactions

Designed to help researchers and practitioners explore, compare, and build state-of-the-art unified multimodal systems.

Awesome Papers & Datasets

Text-and-Image Unified Models

Figure 2: Classification of Unified Multimodal Understanding and Generation Models. The models are divided into three main categories based on their backbone architecture: Diffusion, MLLM (AR), and MLLM (AR + Diffusion). Each category is further subdivided according to the encoding strategy employed, including Pixel Encoding, Semantic Encoding, Learnable Query Encoding, and Hybrid Encoding. We illustrate the architectural variations within these categories and their corresponding encoder-decoder configurations.

Diffusion

MLLM AR

b-1: Pixel Encoding
NameTitleVenueDateCodeDemo
Emu3.5Emu3.5: Native Multimodal Models are World Learners GitHub Repo starsarXiv2025/10/30Github-
Uni-XUni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models GitHub Repo starsICLR2025/09/29Github-
OneCatOneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation GitHub Repo starsarXiv2025/09/03Github-
SelftokSelftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning GitHub Repo starsarXiv2025/05/12Github-
TokLIPTokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation GitHub Repo starsarXiv2025/05/08Github-
HarmonHarmonizing Visual Representations for Unified Multimodal Understanding and Generation GitHub Repo starsarXiv2025/03/27GithubDemo
UGenUGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary LearningarXiv2025/03/27--
SynerGen-VLSynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingarXiv2024/12/12--
LiquidLiquid: Language Models are Scalable and Unified Multi-modal Generators GitHub Repo starsarXiv2024/12/05GithubDemo
OrthusOrthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads GitHub Repo starsarXiv2024/11/28Github-
MMARMMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic ModelingarXiv2024/10/14--
Emu3Emu3: Next-Token Prediction is All You Need GitHub Repo starsarXiv2024/09/27GithubDemo
ANOLEANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation GitHub Repo starsarXiv2024/07/08Github-
ChameleonChameleon: Mixed-Modal Early-Fusion Foundation Models GitHub Repo starsarXiv2024/05/16Github-
LWMWorld Model on Million-Length Video And Language With Blockwise RingAttention GitHub Repo starsICLR2024/02/13Github-
b-2: Semantic Encoding
NameTitleVenueDateCodeDemo
MammothModa2MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and GenerationGitHub Repo starsarXiv2025/11/23GitHub-
Ming-UniVisionMing-UniVision: Joint Image Understanding and Generation with a Unified Continuous TokenizerGitHub Repo starsarXiv2025/10/08GitHub-
Bifrost-1Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP LatentsGitHub Repo starsarXiv2025/08/08GitHub-
Qwen-ImageQwen-Image Technical ReportGitHub Repo starsarXiv2025/08/04GitHubDemo
X-OmniX-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great AgainGitHub Repo starsarXiv2025/07/29GitHubDemo
Ovis-U1Ovis-U1 Technical ReportGitHub Repo starsarXiv2025/06/28GitHubDemo
UniCode$^2$UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and GenerationarXiv2025/06/20--
OmniGen2OmniGen2: Exploration to Advanced Multimodal Generation GitHub Repo starsarXiv2025/06/18GithubDemo
TarVision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations GitHub Repo starsarXiv2025/06/18GithubDemo
UniForkUniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation GitHub Repo starsarXiv2025/06/17Github-
UniWorldUniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation GitHub Repo starsarXiv2025/06/03Github-
PiscesPisces: An Auto-regressive Foundation Model for Image Understanding and Generation GitHub Repo starsarXiv2025/06/10Github-
DualTokenDualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies GitHub Repo starsarXiv2025/03/18Github-
UniTokUniTok: A Unified Tokenizer for Visual Generation and Understanding GitHub Repo starsarXiv2025/02/27GithubDemo
QLIPQLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation GitHub Repo starsarXiv2025/02/05Github-
MetaMorphMetaMorph: Multimodal Understanding and Generation via Instruction Tuning GitHub Repo starsarXiv2024/12/18Github-
ILLUMEILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhancearXiv2024/12/09--
PUMAPUMA: Empowering Unified MLLM with Multi-granular Visual Generation GitHub Repo starsarXiv2024/10/17Github-
VILA-UVILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation GitHub Repo starsICLR2024/09/06GithubDemo
Mini-GeminiMini-Gemini: Mining the Potential of Multi-modality Vision Language Models GitHub Repo starsarXiv2024/03/27GithubDemo
MM-InterleavedMM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer GitHub Repo starsarXiv2024/01/18Github-
VL-GPTVL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation GitHub Repo starsarXiv2023/12/14Github-
Emu2Generative Multimodal Models are In-Context Learners GitHub Repo starsCVPR2023/12/10GithubDemo
DreamLLMDreamLLM: Synergistic Multimodal Comprehension and Creation GitHub Repo starsICLR2023/09/20Github-
LaVITUnified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization GitHub Repo starsICLR2023/09/09Github-
EmuEmu: Generative Pretraining in Multimodality GitHub Repo starsICLR2023/07/11GithubDemo
b-3: Learnable Query Encoding
b-4: Hybrid Encoding (Pseduo)
b-5: Hybrid Encoding (Joint)

MLLM AR-Diffusion

c-1: Pxiel Encoding
c-2: Hybrid Encoding (Pseduo)

Any-to-Any Multimodal models

Benchmark for Evaluation

Benchmarks on Understanding Tasks

Benchmarks on Image Generation Tasks

NamePaperVenueDateCode
GenExamGenExam: A Multidisciplinary Text-to-Image Exam Stararxiv2025/09/18Github
CVTGTextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes Stararxiv2025/08/05Github
OneIG-BenchOneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation Stararxiv2025/06/26Github
ComplexBench-EditComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies Stararxiv2025/06/15Github
EditInspectorEditInspector: A Benchmark for Evaluation of Text-Guided Image Editsarxiv2025/06/11-
ByteMorph-BenchByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions Stararxiv2025/06/03Github
RefEdit-BenchRefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions Stararxiv2025/06/03Github
WISEWISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Stararxiv2025/05/27Github
ImgEdit-BenchImgEdit: A Unified Image Editing Dataset and Benchmark Stararxiv2025/05/26Github
MMIG-BenchMMGen-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Stararxiv2025/05/26Github
KRIS-BenchKRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models StarNeurIPS 20252025/05/22Github
CompBenchCompBench: Benchmarking Complex Instruction-guided Image Editingarxiv2025/05/18-
WorldGenBenchWorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generationarxiv2025/05/02HuggingFace
GEdit-BenchStep1X-Edit: A Practical Framework for General Image Editing StararXiv2025/04/28Github
DreamBench++DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation StarICLR2025/03/09Github
T2I-CompBench++T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation StarTPAMI2025/03/08Github
IE-BenchIE-Bench: Advancing the Measurement of Text-Driven Image Editing for Human Perception AlignmentarXiv2025/01/17-
AnyEditAnyEdit: Mastering Unified High-Quality Image Editing for Any Idea StarCVPR2024/11/24Github
I2EBenchI2EBench: A Comprehensive Benchmark for Instruction-based Image Editing StarNeurIPS2024/08/26Github
ConceptMixConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty StarNeurIPS2024/08/26Github
GenAI-BenchGenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation StarCVPR2024/06/19Github
Commonsense-T2ICommonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? StarCOLM2024/06/11Github
HQ-EditHQ-Edit: A High-Quality Dataset for Instruction-based Image Editing StarICLR2024/04/15Github
VQAScoreEvaluating Text-to-Visual Generation with Image-to-Text Generation StarECCV2024/04/01Github
FlashEvalFlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models StarCVPR2024/03/25Github
DPG-BenchELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment Stararxiv2024/03/08Github
Reason-EditSmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models StarCVPR2023/12/11Github
Emu EditEmu Edit: Precise Image Editing via Recognition and Generation TasksCVPR2023/11/16HuggingFace
HEIMHolistic Evaluation of Text-To-Image Models StarNeurIPS2023/11/07Github
DSG-1kDavidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation StarICLR2023/10/27Github
GenEvalGenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment StarNeurIPS2023/10/17Github
EditValEditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods StararXiv2023/10/03Github
T2I-CompBenchT2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation StarNeurIPS2023/07/12Github
DreamSimDreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data StarNeurIPS2023/06/15Github
MagicBrushMagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing StarNeurIPS2023/06/16Github
MultiGen-20MUniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild StarNeurIPS2023/05/18Github
HRS-BenchHRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models StarICCV2023/04/11Github
TIFATIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering StarICCV2023/03/21Github
EditBenchImagen Editor and EditBench: Advancing and Evaluating Text-Guided Image InpaintingCVPR2022/12/13ProjectPage
PartiPromptsScaling Autoregressive Models for Content-Rich Text-to-Image Generation StarTMLR2022/06/22Github
DrawBenchPhotorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingNeurIPS2022/05/23ProjectPage
PaintSkillsDALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models StarICCV2022/02/08Github

Benchmarks on Interleaved / Compositional / Other Tasks

Dataset

Multimodal Understanding

Text-to-Image

DatasetSamplesPaperVenueDate
FLUX-Reason-6M6MFLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive BenchmarkarXiv2025/09/11
Echo-4o-Image106KEcho-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image GenerationarXiv2025/08/13
Poster100K100KPosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified FrameworkarXiv2025/06/12
Text-Render-2M2MPosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified FrameworkarXiv2025/06/12
ShareGPT-4o-Image45KShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image GenerationarXiv2025/06/22
BLIP3o-60k60KBLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and DatasetarXiv2025/05/14
TextAtlas5M5MTextAtlas5M: A Large-scale Dataset for Dense Text Image GenerationarXiv2025/02/11
EliGen TrainSet500KEliGen: Entity-Level Controlled Image Generation with Regional AttentionarXiv2025/01/02
PD12M12MPublic domain 12m: A highly aesthetic image-text dataset with novel governance mechanismsarXiv2024/10/30
SFHQ-T2I122K--2024/10/06
text-to-image-2M2M--2024/09/13
DenseFusion1MDensefusion-1m: Merging vision experts for comprehensive multimodal perceptionNeurIPS2024/07/11
Megalith10M--2024/07/01
PixelProse16MFrom pixels to prose: A large dataset of dense image captionsarXiv2024/06/14
DOCCI15KDOCCI: Descriptions of Connected and Contrasting ImagesECCV2024/04/30
CosmicMan-HQ 1.06MCosmicman: A text-to-image foundation model for humansCVPR2024/04/01
AnyWord-3M3MAnytext: Multilingual visual text generation and editingICLR2023/11/06
JourneyDB4MJourneyDB: A Benchmark for Generative Image UnderstandingNeurIPS2023/07/03
RenderedText12M--2023/06/30
Mario-10M10MTextdiffuser: Diffusion models as text paintersNeurIPS2023/05/18
SAM11MSegment AnythingICCV2023/04/05
LAION-Aesthetics120MLaion-5b: An open large-scale dataset for training next generation image-text modelsNeurIPS2022/08/16
CC-12M12MConceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual conceptsCVPR2021/02/17

Image Editing

DatasetSamplesPaperVenueDate
Pico-Banana-400K400KPico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingarXiv2025/10/22
X2Edit3.7MX2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation LearningarXiv2025/08/11
ShareGPT-4o-Image (Editing)46KShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image GenerationarXiv2025/06/22
ByteMorph-6M6MByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid MotionsarXiv2025/06/03
ImgEdit1.2MImgEdit: A Unified Image Editing Dataset and BenchmarkarXiv2025/05/26
RefEdit18KRefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model for Referring ExpressionarXiv2025/04/03
AnyEdit2.5MAnyedit: Mastering unified high-quality image editing for any ideaCVPR2024/11/24
OmniEdit1.2MOmniedit: Building image editing generalist models through specialist supervisionICLR2024/11/11
PromptFix1MPromptFix: You Prompt and We Fix the PhotoNeurIPS2024/09/19
UltraEdit4MUltraedit: Instruction-based fine-grained image editing at scaleNeurIPS2024/07/07
EditWorld8.6KEditWorld: Simulating World Dynamics for Instruction-Following Image EditingarXiv2024/06/23
SEED-Data-Edit3.7MSeed-data-edit technical report: A hybrid dataset for instructional image editingarXiv2024/05/07
HQ-Edit197KHq-edit: A high-quality dataset for instruction-based image editingarXiv2024/04/15
HIVE1.1MHIVE: Harnessing Human Feedback for Instructional Visual EditingarXiv2023/07/08
Magicbrush10KMagicbrush: A manually annotated dataset for instruction-guided image editingNeurIPS2023/06/16
InstructP2P313KInstructpix2pix: Learning to follow image editing instructionsCVPR2022/11/17

Interleaved Image-Text

Other Text-Image-to-Image

Applications and Opportunities

Citation

If you find this repo is helpful for your research, please cite our paper:

@article{zhang2025unified,
  title={Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities},
  author={Zhang, Xinjie and Guo, Jintao and Zhao, Shanshan and Fu, Minghao and Duan, Lunhao and Hu, Jiakui and Chng, Yong Xien and Wang, Guo-Hua and Chen, Qing-Guo and Xu, Zhao and Luo, Weihua and Zhang, Kaifu},
  journal={arXiv preprint arXiv:2505.02567},
  year={2025}
}
multimodal-large-language-models
multimodal-models
text-to-image-generation
unified-multimodal-models
vision-language-model

Contributors

suikong-zss

16 commits

sshan-zhao

10 commits

SxJyJay

2 commits

rollingsu

1 commits