friedrichor/Awesome-Multimodal-Papers

A curated list of awesome Multimodal studies.

346

99 commits

updated Aug 4, 2026

See the code

README

Awesome-Multimodal-Papers

A curated list of awesome Multimodal studies.

Contribution

If you have published a high-quality paper or come across one that you think is valuable, feel free to contribute! To submit a paper, please open an issue and include the following information in the specified format:

Submission Format
{
    "title": paper title,
    "url": paper URL,
    "venue": the venue where the paper was published, such as ICML 2025, CVPR 2025 or arXiv,
    "category": one or more relevant categories from our directory, or feel free to propose a new, more suitable category,
    "code": [Optional] code URL,
    "project_page": [Optional] project page URL,
    "dataset": [Optional] HuggingFace Dataset URL,
    "collections": [Optional] HuggingFace Collections URL
}

Foundation Model (Textual and Multimodal)

Visual Understanding

TitleVenueDateCodeSupplement
VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionACM MM 20262026-07-14Star-
TableDART: Dynamic Adaptive Multi-Modal Routing for Table UnderstandingICLR 20262025-09-18Star-
Adaptive MLP Pruning for Large Vision Transformers (AMP)arXiv2026-03-10StarModel
MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular LearningCVPR 20262026-02-23Star-
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers (DGMR)arXiv2025-06-10StarModel
Learning Compact Vision Tokens for Efficient Large Multimodal Models (LLaVA-STF)arXiv2025-06-08StarModel
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal ModelsarXiv2025-04-14StarProject Page
Collections
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression (Baidu)arXiv2025-03-27Star-
M-LLM Based Video Frame Selection for Efficient Video Understanding (CMU)arXiv2025-02-27--
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMsICLR 20252025-02-24Star-
Qwen2.5 VL-2025-01-26StarProject Page
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context ModelingarXiv2025-01-21Star-
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding (Spatial-Temporal Compression)arXiv2025-01-14Star-
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video UnderstandingarXiv2025-01-09Star-
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and VideosarXiv2025-01-07StarProject Page
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models (Video Token Compression)arXiv2024-12-30StarProject Page
Apollo: An Exploration of Video Understanding in Large Multimodal Models (Exploration) (Meta)arXiv2024-12-13StarProject Page
CompCap: Improving Multimodal Large Language Models with Composite Captions (Meta)arXiv2024-12-09--
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling (InternVL 2.5)arXiv2024-12-06StarProject Page
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMsarXiv2024-10-21-Project Page
[Model, Dataset] Personalized Visual Instruction Tuning (PVIT, PVIT-3M)arXiv2024-10-09StarDataset
Video Instruction Tuning With Synthetic Data (LLaVA-Video, LLaVA-NeXT Series)arXiv2024-10-03StarProject Page
Dataset
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsarXiv2024-09-26-Project Page
Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model (MGLMM, Alibaba)arXiv2024-09-20StarProject Page
POINTS: Improving Your Vision-language Model with Affordable Strategies (WeChat)arXiv2024-09-07Star-
xGen-MM (BLIP-3): A Family of Open Large Multimodal ModelsarXiv2024-08-16StarProject Page
Collections
LLaVA-OneVision: Easy Visual Task Transfer (LLaVA-NeXT Series)arXiv2024-08-06StarProject Page
Dataset
Tarsier: Recipes for Training and Evaluating Large Video Description Models (Tarsier, Dream1k, by ByteDance)arXiv2024-07-30StarDataset
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and OutputarXiv2024-07-03Star-
TokenPacker: Efficient Visual Projector for Multimodal LLMarXiv2024-07-02Star-
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs (Cambrian, Data Rationing)arXiv2024-06-24StarProject Page
Dataset
Long Context Transfer from Language to Vision (LongVA, by Ziwei Liu, Chunyuan Li)arXiv2024-06-24StarProject Page
Generative Visual Instruction TuningarXiv2024-06-17Star-
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video UnderstandingarXiv2024-06-13StarCollections
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities (Apple)arXiv2024-06-13StarProject Page
Wechat
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMsarXiv2024-06-11Star-
Wings: Learning Multimodal LLMs without Text-only ForgettingarXiv2024-06-05--
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment (MIVPG)arXiv2024-06-05--
PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLMarXiv2024-06-05Star-
OLIVE: Object Level In-Context Visual EmbeddingsACL 20242024-06-02Star-
X-VILA: Cross-Modality Alignment for Large Language Model (by NVIDIA)arXiv2024-05-29-Wechat
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal ModelsarXiv2024-05-24Star-
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language ModelsarXiv2024-05-24--
LOVA3: Learning to Visual Question Answering, Asking and AssessmentarXiv2024-05-23Star-
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityarXiv2024-05-23StarProject Page
Chameleon: Mixed-Modal Early-Fusion Foundation Models (Meta)arXiv2024-05-16StarBlog
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-ExpertsarXiv2024-05-09StarProject Page
Dataset
Wechat
ImageInWords: Unlocking Hyper-Detailed Image Descriptions (Google)arXiv2024-05-05StarProject Page
Dataset
What matters when building vision-language models? (Idefics2)arXiv2024-05-03-
Collections
MANTIS: Interleaved Multi-Image Instruction TuningarXiv2024-05-02StarProject Page
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMsCVPR 2024 Workshop2024-04-23-Wechat
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language ModelsarXiv2024-04-19StarProject Page
Dataset
MoVA: Adapting Mixture of Vision Experts to Multimodal ContextarXiv2024-04-19Star-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language ModelsarXiv2024-04-18-Project Page
Project Page
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation? (LaDiC)NAACL 20242024-04-16Star-
AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception (AesExpert, AesMMIT Dataset)arXiv2024-04-15Star-
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models (Ferret-v2)arXiv2024-04-11--
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies (MiniCPM series)arXiv2024-04-09Star
Star
Blog
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (Ferret-UI)arXiv2024-04-08--
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingCVPR 20242024-04-08StarProject Page
Koala: Key frame-conditioned long video-LLMCVPR 20242024-04-05StarProject Page
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual TokensarXiv2024-04-04StarProject Page
LongVLM: Efficient Long Video Understanding via Large Language ModelsarXiv2024-04-04Star-
InternVideo2: Scaling Foundation Models for Multimodal Video UnderstandingECCV 20242024-03-22Star-
VideoAgent: Long-form Video Understanding with Large Language Model as Agent (key frame)arXiv2024-03-15--
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training (Apple)arXiv2024-03-14--
UniCode: Learning a Unified Codebook for Multimodal Large Language ModelsarXiv2024-03-14--
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextarXiv2024-03-08-Project Page
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language ModelsarXiv2023-03-05Star-
RegionGPT: Towards Region Understanding Vision Language ModelCVPR 20242024-03-04-Project Page
All in an Aggregated Image for In-Image LearningarXiv2024-02-28Star-
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersCVPR 20242024-02-27StarProject Page
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different LanguagesarXiv2024-02-25--
LLMBind: A Unified Modality-Task Integration FrameworkarXiv2024-02-22--
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model (ALLaVA)arXiv2024-02-18StarDemo Page
Dataset
MobileVLM V2: Faster and Stronger Baseline for Vision Language ModelarXiv2024-02-06Star-
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile DevicesarXiv2023-12-28Star-
Gemini: A Family of Highly Capable Multimodal ModelsarXiv2023-12-19-Project Page
Osprey: Pixel Understanding with Visual Instruction TuningCVPR 20242023-12-15Star-
VILA: On Pre-training for Visual Language Models (NVIDIA, MIT)CVPR 20242023-12-12Star-
Vary: Scaling up the Vision Vocabulary for Large Vision-Language ModelsarXiv2023-12-11StarProject Page
Prompt Highlighter: Interactive Control for Multi-Modal LLMsCVPR 20242023-12-07StarProject Page
PixelLM: Pixel Reasoning with Large Multimodal ModelCVPR 20242023-12-04StarProject Page
APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsEMNLP 20232023-12-04StarProject Page
LLaMA-VID: An Image is Worth 2 Tokens in Large Language ModelsarXiv2023-11-28StarProject Page
Dataset
PG-Video-LLaVA: Pixel Grounding Large Video-Language ModelsarXiv2023-11-22Star-
ShareGPT4V: Improving Large Multi-Modal Models with Better CaptionsarXiv2023-11-21StarProject Page
LION : Empowering Multimodal Large Language Model with Dual-Level Visual KnowledgeCVPR 20242023-11-20StarProject Page
Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionarXiv2023-11-16Star-
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality CollaborationarXiv2023-11-07Star-
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learningarXiv2023-10-14StarProject Page
Ferret: Refer and Ground Anything Anywhere at Any Granularity (Ferret)ICLR 20242023-10-11Star-
Improved Baselines with Visual Instruction Tuning (LLaVA-1.5)arXiv2023-10-05StarProject Page
Aligning Large Multimodal Models with Factually Augmented RLHF (LLaVA-RLHF, MMHal-Bench (hallucination))arXiv2023-09-25StarProject Page
Dataset
MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningICLR 20242023-09-14Star-
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and BeyondarXiv2023-08-24StarProject Page
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages (VisCPM-Chat/Paint)ICLR 20242023-08-23Star-
SVIT: Scaling up Visual Instruction TuningarXiv2023-07-09StarDataset
Kosmos-2: Grounding Multimodal Large Language Models to the World (Kosmos-2, GrIT Dataset)arXiv2023-06-26StarDemo
Dataset
M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction TuningarXiv2023-06-07-Project Page
Dataset
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningNeurIPS 20232023-05-11Star-
MultiModal-GPT: A Vision and Language Model for Dialogue with HumansarXiv2023-05-08Star-
VPGTrans: Transfer Visual Prompt Generator across LLMsNeurIPS 20232023-05-02StarProject Page
mPLUG-Owl: Modularization Empowers Large Language Models with MultimodalityarXiv2023-04-27Star-
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsICLR 20242023-04-20StarProject Page
Visual Instruction Tuning (LLaVA)NeurIPS 20232023-04-17StarProject Page
Dataset
Language Is Not All You Need: Aligning Perception with Language Models (Kosmos-1)NeurIPS 20232023-02-27Star-
Multimodal Chain-of-Thought Reasoning in Language ModelsarXiv2023-02-02Star-
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsICML 20232023-01-30Star-
Flamingo: a Visual Language Model for Few-Shot LearningNeurIPS 20222022-04-29Star-

Omni Understanding

TitleVenueDateCodeSupplement
Ming-Omni: A Unified Multimodal Model for Perception and Generation (Ant Group)arXiv2025-06-11StarProject Page
Qwen2.5-Omni Technical ReportarXiv2025-03-26StarProject Page
Collections
PAVE: Patching and Adapting Video Large Language ModelsCVPR 20252025-03-25StarDataset
Baichuan-Omni-1.5 Technical ReportarXiv2025-01-26Star-
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement LearningarXiv2025-03-07StarModel
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferencearXiv2025-02-25StarProject Page
Collections
Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment (THU, Tencent Hunyuan, NTU S-Lab)arXiv2025-02-06StarProject Page
[Benchmark] WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs (Xiaohongshu, SJTU)arXiv2025-02-06StarProject Page
Dataset
Align Anything: Training All-Modality Models to Follow Instructions with Language FeedbackarXiv2024-12-20StarProject Page
Dataset
[Survey] From Specific-MLLM to Omni-MLLM: A Survey about the MLLMs alligned with Multi-Modality (HIT, Peng Cheng Lab)arXiv2024-12-16--
OMCAT: Omni Context Aware Transformer (OCTAV, OMCAT) (NVIDIA)arXiv2024-10-15-Project Page
Baichuan-Omni Technical ReportarXiv2024-10-11Star-
[Benchmark] OmniBench: Towards The Future of Universal Omni-Language ModelsarXiv2024-09-23StarProject Page
Dataset
OmniBind: Large-scale Omni Multimodal Representation via Binding SpacesarXiv2024-07-16StarProject Page
Explore the Limits of Omni-modal Pretraining at Scale (MiCo)arXiv2024-06-13StarProject Page
ViT-Lens: Towards Omni-modal Representations (TencentARC)CVPR 20242023-08-20StarProject Page
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetNeurIPS 20232023-05-29Star-
ImageBind: One Embedding Space To Bind Them AllCVPR 20232023-05-09StarProject Page
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and DatasetTPAMI 20242023-04-17StarProject Page

Unified Understanding and Generation

TitleVenueDateCodeSupplement
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal VelocitiesarXiv2025-05-26-Project Page
Wechat
MMaDA: Multimodal Large Diffusion Language Models (ByteDance Seed)arXiv2025-05-21StarDemo
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation (Apple)arXiv2025-05-20--
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation (ByteDance Seed)arXiv2025-05-08--
[Survey] Unified Multimodal Understanding and Generation Models: Advances, Challenges, and OpportunitiesarXiv2025-05-05Star-
Unified Reward Model for Multimodal Understanding and Generation (UnifiedReward) (Fudan, Shanghai AI Lab)arXiv2025-03-07StarProject Page
Dataset
UniTok: A Unified Tokenizer for Visual Generation and Understanding (ByteDance)arXiv2025-02-27StarProject Page
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling (by deepseek)arXiv2025-01-29Star-
LlamaFusion: Adapting Pretrained Language Models for Multimodal Generation (Meta)arXiv2024-12-19--
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingarXiv2024-12-12coming soon-
ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhancearXiv2024-12-09-Wechat
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation (TencentARC)arXiv2024-12-05Star-
Liquid: Language Models are Scalable Multi-modal Generators (Bytedance)arXiv2024-12-05StararXiv
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation (ByteDance)arXiv2024-12-04StarProject Page
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific HeadsICML 20252024-11-28Star-
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation (by deepseek)-2024-10-17Star-
Emu3: Next-Token Prediction is All You NeedarXiv2024-09-27StarProject Page
Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationarXiv2024-08-22StarProject Page
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal ModelICLR 2025 Oral2024-08-20Star-
An Image is Worth 32 Tokens for Reconstruction and Generation (TiTok, by ByteDance)arXiv2024-06-11StarProject Page
Mini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsarXiv2024-05-27StarProject Page
Collections
VITRON: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing-2024-04-25StarProject Page
YouTube
Wechat
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and GenerationarXiv2024-04-22Star-
AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingarXiv2024-02-19StarProject Page
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationarXiv2024-02-05StarProject Page
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and ActionarXiv2023-12-28StarProject Page
Generative Multimodal Models are In-Context Learners (Emu2)CVPR 20242023-12-20StarProject Page
CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any GenerationarXiv2023-11-30StarProject Page
LLMGA: Multimodal Large Language Model based Generation AssistantarXiv2023-11-27StarProject Page
VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and GenerationarXiv2023-12-14Star-
Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsICLR 20242023-10-04StarProject Page
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative VokensarXiv2023-10-03StarProject Page
DreamLLM: Synergistic Multimodal Comprehension and CreationICLR 20242023-09-20StarProject Page
NExT-GPT: Any-to-Any Multimodal LLMarXiv2023-09-11StarProject Page
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization (LaVIT)ICLR 20242023-09-09Star-
Planting a SEED of Vision in Large Language ModelICLR 20242023-07-16StarProject Page
Generative Pretraining in Multimodality (Emu1)ICLR 20242023-07-11Star-
Generating Images with Multimodal Language Models (GILL)NeurIPS 20232023-05-26StarProject Page
Any-to-Any Generation via Composable Diffusion (CoDi-1)NeurIPS 20232023-05-19StarProject Page
Grounding Language Models to Images for Multimodal Inputs and Outputs (FROMAGe)ICML 20232023-01-31StarProject Page

Coding / GUI

Vibe Coding / GUI / UI Design

TitleVenueDateCodeSupplement
[Model, Benchmark] UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation (THU, Zhipu)arXiv2025-11-14StarProject Page
[Agent] Computer-Use Agents as Judges for Generative User Interface (AUI) (Oxon, NUS Show Lab, Microsoft)arXiv2025-11-19StarProject Page
[Agent, Benchmark] Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development (TDDev)arXiv2025-09-29Star-
[Model] UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement LearningarXiv2025-09-02StarProject Page
Showcase
[Agent, Benchmark] FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI AgentsarXiv2025-08-12--
[Agent] ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal AgentsarXiv2025-07-30StarDemo
[Benchmark] ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation EvaluationarXiv2025-07-07StarProject Page
Dataset
[Benchmark] DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code GenerationarXiv2025-06-06StarDataset
[Model, Benchmark] WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from ScratcharXiv2025-05-06StarProject Page
Dataset
Dataset
[Benchmark] GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent (GTArena)arXiv2024-12-24StarDataset
[Benchmark] Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive PrototypingarXiv2024-11-05StarProject Page
Dataset
[Agent, Benchmark] Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design PrototypingNAACL 20252024-10-21StarProject Page
Dataset
[Dataset] WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsWWW 20252024-04-09StarProject Page
Dataset
[Benchmark] Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End EngineeringNAACL 20252024-03-05StarProject Page
Dataset

Diffusion MLLM

Multimodal Embedding/Retrieval

TitleVenueDateCodeSupplement
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought ReasoningEMNLP 20252025-09-25StarProject Page
Dataset
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual RetrievalarXiv2025-06-23-Project Page
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval (UNITE)arXiv2025-05-26StarProject Page
Collections
Wechat
[Benchmark] MIEB: Massive Image Embedding BenchmarkarXiv2025-04-14StarDemo
[Data, Model, Benchmark] IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal RetrievalarXiv2025-04-01Star-
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive LearningarXiv2025-03-04StarCollections
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval (CoT)arXiv2025-02-28--
[Model, Dataset] Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval (FDCA, FineCVR-1M)ICLR 20252025-02-26StarProject Page
[Benchmark, Model] MomentSeeker: A Comprehensive Benchmark and A Strong Baseline For Moment Retrieval Within Long Videos (MomentSeeker, V-Embedder) (Gaoling)arXiv2025-02-18StarDataset
[Data, Model, Benchmark] Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval (Vis-IR task, VIRA, UniSE, MVRB)arXiv2025-02-17--
[Data, Model] Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval (FineCVR-1M, FDCA)ICLR 20252025-01-23StarProject Page
[Benchmark] CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval (CaReBench)CVPR 20252024-12-31StarProject Page
Dataset
MINIMA: Modality Invariant Image MatchingCVPR 20252024-12-27StarDemo
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs (Tongyi Lab)CVPR 20252024-12-22StarCollections
✨ [Dataset, Model] MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval (MagaPairs, BGE-VL)arXiv2024-12-19StarDataset
Wechat
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval (OSrCIR) (CoT)CVPR 20252024-12-15Star-
LamRA: Large Multimodal Model as Your Advanced Retrieval AssistantarXiv2024-12-02StarProject Page
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs (NVIDIA)ICLR 2025 Poster2024-11-04-Model
OMCAT: Omni Context Aware TransformerarXiv2024-10-15-Project Page
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks (TIGER Lab)arXiv2024-10-07StarProject Page
Dataset
Dataset
Collections
E5-V: Universal Embeddings with Multimodal Large Language ModelsarXiv2024-07-17Star-
OmniBind: Large-scale Omni Multimodal Representation via Binding SpacesarXiv2024-07-16StarProject Page
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models (NVIDIA)ICLR 20252024-05-27-Collections
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions (Google DeepMind)ICML 2024 Oral2024-03-28StarProject Page
DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation ModelsNAACL2024-04-07--
Composed Video Retrieval via Enriched Context and Discriminative EmbeddingsCVPR 20242024-03-25Star-
UniIR: Training and Benchmarking Universal Multimodal Information Retrievers (TIGER Lab)ECCV 20242023-11-28StarProject Page
Dataset
CoVR-2: Automatic Data Construction for Composed Video Retrieval&CoVR: Learning Composed Video Retrieval from Web Video CaptionsTPAMI 2024 & AAAI 20242023-08-23StarProject Page
Dataset
Dataset
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetNeurIPS 20232023-05-29Star-
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and DatasetTPAMI20242023-04-17StarProject Page
✨ (QB-Norm) Cross Modal Retrieval with Querybank NormalisationCVPR 20222021-12-23StarProject Page
✨ (DSL) Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax LossarXiv2021-09-09Star-

Image Understanding Benchmark

Video Understanding Benchmark

TitleVenueDateCodeSupplement
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsACL 20262026-03StarProject Page
Dataset
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic VideosACL 2025 Main2025-05-26StarProject Page
Dataset
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation ModelsarXiv2024-10-30StarDataset
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark (AuroraCap, VDC)arXiv2024-10-24StarProject Page
Dataset
Dataset
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video ModelsarXiv2024-10-14StarProject Page
Dataset
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingNeurIPS 20242024-09-26StarProject Page
Dataset
Dataset
Tarsier: Recipes for Training and Evaluating Large Video Description Models (Tarsier, Dream1k) (ByteDance)arXiv2024-07-30StarDataset
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning (VideoVista)arXiv2024-06-17StarProject Page
Dataset
VELOCITI: Can Video-Language Models Bind Semantic Concepts through Time?arXiv2024-06-16StarProject Page
Dataset
MLVU: A Comprehensive Benchmark for Multi-Task Long Video UnderstandingarXiv2024-06-06StarDataset
Dataset
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis (Video-MME)arXiv2024-05-31StarProject Page
Dataset
TempCompass: Do Video LLMs Really Understand Videos?arXiv2024-03-01StarProject Page
Dataset
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark (MVBench, VideoChat2)CVPR 2024 Highlight2023-11-28StarDataset
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language UnderstandingNeurIPS 20232023-08-17StarProject Page
Dataset
Perception Test: A Diagnostic Benchmark for Multimodal Video Models (Perception Test, by Google DeepMind)NeurIPS 20232023-05-23StarProject Page
Dataset

Audio

Multimodal Dialogue

Multimodal Learning

Image Generation

TitleVenueDateCodeSupplement
ImageBench V1: Text-to-Image Benchmark with VLM Judges (live benchmark; 40+ models, 192 prompts, all outputs published)WebsiteLive-Project Page Methodology
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and EditingarXiv2025-03-13Star-
OmniGen: Unified Image GenerationarXiv2024-09-17StarProject Page
Demo
Wechat
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens (by Kaiming He, DeepMind, MIT)arXiv2024-10-17--
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (Lumina-T2X, Flag-DiT) (Text2Any)arXiv2024-05-09StarYouTube
Wechat
FreeU: Free Lunch in Diffusion U-Net (FreeU, by Ziwei Liu)CVPR 2024 Oral2023-09-20StarProject Page
YouTube
Demo
Lazy Diffusion Transformer for Interactive Image EditingarXiv2024-04-18-Project Page
Salient Object-Aware Background Generation using Text-Guided Diffusion ModelsCVPR 2024 Workshop2024-04-15Star-
HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingarXiv2024-04-15StarProject Page
Dataset
Demo
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark (UNIAA-LLaVA, UNIAA-Bench)arXiv2024-04-15--
PMG: Personalized Multimodal Generation with Large Language ModelsWWW 20242024-04-07--
Identity Decoupling for Multi-Subject Personalization of Text-to-Image ModelsarXiv2024-04-05StarProject Page
Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsCVPR 20242024-04-05--
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (VAR)arXiv2024-04-03StarProject Page
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation (HuaWei, Enze Xie)arXiv2024-03-07StarProject Page
Multi-LoRA Composition for Image GenerationarXiv2024-02-26StarProject Page
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models (HuaWei, Enze Xie)arXiv2024-01-10StarProject Page
Brush Your Text: Synthesize Any Scene Text on Images via Diffusion ModelAAAI 20242023-12-19Star-
SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models (Tencent Xintao Wang)arXiv2023-12-11StarProject Page
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction FollowingarXiv2023-12-11StarProject Page
Emu Edit: Precise Image Editing via Recognition and Generation TasksarXiv2023-11-16-Project Page
BeautifulPrompt: Towards Automatic Prompt Engineering for Text-to-Image SynthesisEMNLP 20232023-11-12Starzhihu
AnyText: Multilingual Visual Text Generation And EditingICLR 20242023-11-06Star-
EasyGen: Easing Multimodal Generation with a Bidirectional Conditional Diffusion Model and LLMsarXiv2023-10-13Star-
Mini-DALLE3: Interactive Text to Image by Prompting Large Language ModelsarXiv2023-10-11StarProject Page
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis (HuaWei, Enze Xie)ICLR 2024 Spotlight2023-09-30StarProject Page
Dataset
Usage Diffusers
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion ModelsarXiv2023-08-13StarProject Page
Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsarXiv2023-10-04StarProject Page
Improving Image Generation with Better Captions (DALL-E 3)OpenAI2023--
Scaling up GANs for Text-to-Image Synthesis (GigaGAN)CVPR 20232023-05-09StarProject Page
Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)ICCV 20232023-02-10Star-
Scalable Diffusion Models with Transformers (DiT)ICCV 20232022-12-19StarProject Page
InstructPix2Pix: Learning to Follow Image Editing InstructionsCVPR 20232022-11-17StarProject Page
All are Worth Words: A ViT Backbone for Diffusion Models (U-ViT, first Diffsuion Transformer) (RUC, Chongxuan Li)CVPR 20232022-09-25Star-
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven GenerationCVPR 20232022-08-25StarProject Page
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding (Imagen)NeurIPS 20222022-05-23StarProject Page
Hierarchical Text-Conditional Image Generation with CLIP Latents (DALL-E 2)OpenAI2022-04-13Star-
High-Resolution Image Synthesis with Latent Diffusion Models (LDM, Stable Diffusion)CVPR 20222021-12-20Star-
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsICML 20222021-12-20Star-
NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtionECCV 20222021-11-24Star-
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsICLR 20222021-08-02StarProject Page
CogView: Mastering Text-to-Image Generation via TransformersNeurIPS 20212021-05-26Star-
Zero-Shot Text-to-Image Generation (DALL-E 1)ICML 20212021-02-24StarProject Page
Taming Transformers for High-Resolution Image Synthesis (VQ-GAN)CVPR 20212020-12-17StarProject Page

Video Generation

TitleVenueDateCodeSupplement
ReCamMaster: Camera-Controlled Generative Rendering from A Single VideoICCV20252025-03-14StarProject Page
Dataset
[Dataset] Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video SpecialistsarXiv2025-02-10StarProject Page
Dataset
Demo
LiFT: Leveraging Human Feedback for Text-to-Video Model AlignmentarXiv2024-12-06StarProject Page
[Dataset] VidGen-1M: A Large-Scale Dataset for Text-to-video GenerationarXiv2024-08-05StarProject Page
Dataset
[Dataset] MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions (Mira)arXiv2024-07-08Video GenerationStar
VIMI: Grounding Video Generation through Multi-modal InstructionarXiv2024-07-08StarProject Page
[Dataset] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationICLR 20252024-07-02StarProject Page
Dataset
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (Lumina-T2X, Flag-DiT) (Text2Any)arXiv2024-05-09StarYouTube
Wechat
StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text (Long Video Generation)arXiv2024-03-21StarProject Page
YouTube
Demo
Wechat
AnyV2V: A Plug-and-Play Framework For Any Video-to-Video Editing TasksarXiv2024-03-21StarProject Page
Demo
Demo Page
Wechat
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation (FRESCO) (NTU, Ziwei Liu)CVPR 20242024-03-19StarProject Page
Latte: Latent Diffusion Transformer for Video Generation (Latte) (NTU, Ziwei Liu)arXiv2024-01-05StarProject Page
FreeInit: Bridging Initialization Gap in Video Diffusion Models (FreeInit) (NTU, Ziwei Liu)arXiv2023-12-12StarProject Page
YouTube
Demo
VideoBooth: Diffusion-based Video Generation with Image Prompts (VideoBooth) (NTU, Ziwei Liu)arXiv2023-12-01StarProject Page
VBench: Comprehensive Benchmark Suite for Video Generative Models [Benchmark] (VBench) (NTU, Ziwei Liu)CVPR 20242023-11-29StarProject Page
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (SVD)arXiv2023-11-25StarProject Page
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction (NTU, Ziwei Liu)ICLR 20242023-10-31StarProject Page
FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling (FreeNoise) (NTU, Ziwei Liu)ICLR 20242023-10-23StarProject Page
LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models (LaVie) (NTU, Ziwei Liu)arXiv2023-09-26StarProject Page

Multimodal Dataset

TitleVenueDateCodeSupplement
[Benchmark & Dataset] VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsarXiv2025-04-21Visual Reasoning Benchmark & Dataset🌐 Homepage
🤗 Benchmark
🤗 Train Data
💻 Train Code
[Dataset] Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video SpecialistsarXiv2025-02-10StarProject Page
Dataset
Demo
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction DataarXiv2024-10-24-Dataset
LVD-2M: A Long-take Video Dataset with Temporally Dense CaptionsNeurIPS 20242024-10-14StarProject Page
VidGen-1M: A Large-Scale Dataset for Text-to-video GenerationarXiv2024-08-05StarProject Page
Dataset
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions (Mira)arXiv2024-07-08Video GenerationStar
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationICLR 20252024-07-02StarProject Page
Dataset
GUIDE: A Guideline-Guided Dataset for Instructional Video ComprehensionIJCAI 20242024-06-26-Project Page
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion TokensarXiv2024-06-17StarCollections
Blog
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and GenerationarXiv2024-06-15Star-
What If We Recaption Billions of Web Images with LLaMA-3? (Recap-DataComp-1B)arXiv2024-06-12StarProject Page
Dataset
[Dataset
TextSquare: Scaling up Text-Centric Visual Instruction TuningarXiv2024-04-19Visual Instruction Tuning-
HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingarXiv2024-04-15Instruction Image EditingStar
Project Page
Dataset
Demo
AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception (AesExpert, AesMMIT Dataset)arXiv2024-04-15Aesthetic Multi-Modality Instruction TuningStar
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality TeachersCVPR 20242024-02-29video-captionStar
Project Page
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language ModelarXiv2024-02-18GPT4V-synthesized DataStar
Demo Page
Dataset
STICKERCONV: Generating Multimodal Empathetic Responses from ScratchACL 2024 Main2024-01-20Multimodal Empathetic DialogueStar
Project Page
Dataset
Wechat
zhihu
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationICLR 20242023-07-13StarDataset
Dataset
SVIT: Scaling up Visual Instruction TuningarXiv2023-07-09Instruction TuningStar
Dataset
Kosmos-2: Grounding Multimodal Large Language Models to the World (Kosmos-2, GrIT Dataset)arXiv2023-06-26Grounded image-text pairsStar
Demo
Dataset
M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction TuningarXiv2023-06-07Instruction TuningProject Page
Dataset
Visual Instruction Tuning (LLaVA)NeurIPS 20232023-04-17Instruction TuningStar
Project Page
Dataset
Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with TextNeurIPS D&B 20232023-04-14Interleaved Image-TextStar
AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content Creation (Wiki)TREC 2023 Workshop2023-04-04StarDataset
Dataset
TikTalk: A Multi-Modal Dialogue Dataset for Real-World ChitchatACM MM 20232023-01-14Multimodal DialogueStar
MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationACL 20232022-11-10Multimodal DialogueStar
LAION-5B: An open large-scale dataset for training next generation image-text modelsNeurIPS 20222022-10-16Image-Text PairsProject Page
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text PairsNeurIPS Workshop 20212021-11-03Image-Text PairsProject Page
MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsACM SIGIR 20212021-07Multimodal DialogueStar
PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text ModelingACL 20212021-07-06Open-domain Multimodal DialogueStar
Image-Chat: Engaging Grounded ConversationsACL 20202018-11-02Multimodal DialogueProject Page

Multimodal Survey

TitleVenueDateSupplementLatest Update
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and ApplicationsTMLR 20262026-05--
Discrete Diffusion in Large Language and Multimodal Models: A SurveyarXiv2025-06-16Star-
A Survey on Bridging VLMs and Synthetic DataOpenReview2025-05-16Star
Github Page
-
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and OpportunitiesarXiv2025-05-05Star-
From Specific-MLLM to Omni-MLLM: A Survey about the MLLMs alligned with Multi-Modality (HIT, Peng Cheng Lab)arXiv2024-12-16-
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding (TikTok)arXiv2024-09-27-
Video Diffusion Models: A SurveyarXiv2024-05-06Star-
Theoretical research on generative diffusion models: an overviewarXiv2024-04-13--
A Review of Multi-Modal Large Language and Vision ModelsarXiv2024-03-28--
The (R)Evolution of Multimodal Large Language Models: A SurveyarXiv2024-02-19--
MM-LLMs: Recent Advances in MultiModal Large Language ModelsarXiv2024-01-24-2024-02-20
Multimodal Large Language Models: A SurveyIEEE BigData 20232023-11-22--
Multimodal Foundation Models: From Specialists to General-Purpose AssistantsCVPR 20232023-09-18--
Understanding Deep Learning-2023--
Large Multimodal Models: Notes on CVPR 2023 TutorialCVPR 20232023-06-26--
A Survey on Multimodal Large Language ModelsarXiv2023-06-23-2024-04-01
Multimodal Deep LearningarXiv2023-01-12--
Diffusion Models: A Comprehensive Survey of Methods and ApplicationsACM Computing Surveys2022-09-02-2024-02-06
Multimodal Learning with Transformers: A SurveyIEEE TPAMI 20232022-01-13-2023-05-10
Multimodal Machine Learning: A Survey and TaxonomyIEEE PAMI 20192017-05-26-2017-08-01
deep-learning
large-multimodal-models
multimodal
multimodal-data
multimodal-deep-learning
multimodal-dialogue
multimodal-large-language-models
multimodal-learning

Contributors

friedrichor

94 commits

mghiasvand1

3 commits

dh7

1 commits

xwy-bit

1 commits

friedrichor/Awesome-Multimodal-Papers

A curated list of awesome Multimodal studies.

346

99 commits

updated Aug 4, 2026

See the code

README

Awesome-Multimodal-Papers

A curated list of awesome Multimodal studies.

Contribution

If you have published a high-quality paper or come across one that you think is valuable, feel free to contribute! To submit a paper, please open an issue and include the following information in the specified format:

Submission Format
{
    "title": paper title,
    "url": paper URL,
    "venue": the venue where the paper was published, such as ICML 2025, CVPR 2025 or arXiv,
    "category": one or more relevant categories from our directory, or feel free to propose a new, more suitable category,
    "code": [Optional] code URL,
    "project_page": [Optional] project page URL,
    "dataset": [Optional] HuggingFace Dataset URL,
    "collections": [Optional] HuggingFace Collections URL
}

Foundation Model (Textual and Multimodal)

Visual Understanding

TitleVenueDateCodeSupplement
VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token CompressionACM MM 20262026-07-14Star-
TableDART: Dynamic Adaptive Multi-Modal Routing for Table UnderstandingICLR 20262025-09-18Star-
Adaptive MLP Pruning for Large Vision Transformers (AMP)arXiv2026-03-10StarModel
MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular LearningCVPR 20262026-02-23Star-
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers (DGMR)arXiv2025-06-10StarModel
Learning Compact Vision Tokens for Efficient Large Multimodal Models (LLaVA-STF)arXiv2025-06-08StarModel
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal ModelsarXiv2025-04-14StarProject Page
Collections
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression (Baidu)arXiv2025-03-27Star-
M-LLM Based Video Frame Selection for Efficient Video Understanding (CMU)arXiv2025-02-27--
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMsICLR 20252025-02-24Star-
Qwen2.5 VL-2025-01-26StarProject Page
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context ModelingarXiv2025-01-21Star-
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding (Spatial-Temporal Compression)arXiv2025-01-14Star-
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video UnderstandingarXiv2025-01-09Star-
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and VideosarXiv2025-01-07StarProject Page
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models (Video Token Compression)arXiv2024-12-30StarProject Page
Apollo: An Exploration of Video Understanding in Large Multimodal Models (Exploration) (Meta)arXiv2024-12-13StarProject Page
CompCap: Improving Multimodal Large Language Models with Composite Captions (Meta)arXiv2024-12-09--
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling (InternVL 2.5)arXiv2024-12-06StarProject Page
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMsarXiv2024-10-21-Project Page
[Model, Dataset] Personalized Visual Instruction Tuning (PVIT, PVIT-3M)arXiv2024-10-09StarDataset
Video Instruction Tuning With Synthetic Data (LLaVA-Video, LLaVA-NeXT Series)arXiv2024-10-03StarProject Page
Dataset
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsarXiv2024-09-26-Project Page
Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model (MGLMM, Alibaba)arXiv2024-09-20StarProject Page
POINTS: Improving Your Vision-language Model with Affordable Strategies (WeChat)arXiv2024-09-07Star-
xGen-MM (BLIP-3): A Family of Open Large Multimodal ModelsarXiv2024-08-16StarProject Page
Collections
LLaVA-OneVision: Easy Visual Task Transfer (LLaVA-NeXT Series)arXiv2024-08-06StarProject Page
Dataset
Tarsier: Recipes for Training and Evaluating Large Video Description Models (Tarsier, Dream1k, by ByteDance)arXiv2024-07-30StarDataset
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and OutputarXiv2024-07-03Star-
TokenPacker: Efficient Visual Projector for Multimodal LLMarXiv2024-07-02Star-
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs (Cambrian, Data Rationing)arXiv2024-06-24StarProject Page
Dataset
Long Context Transfer from Language to Vision (LongVA, by Ziwei Liu, Chunyuan Li)arXiv2024-06-24StarProject Page
Generative Visual Instruction TuningarXiv2024-06-17Star-
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video UnderstandingarXiv2024-06-13StarCollections
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities (Apple)arXiv2024-06-13StarProject Page
Wechat
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMsarXiv2024-06-11Star-
Wings: Learning Multimodal LLMs without Text-only ForgettingarXiv2024-06-05--
Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment (MIVPG)arXiv2024-06-05--
PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLMarXiv2024-06-05Star-
OLIVE: Object Level In-Context Visual EmbeddingsACL 20242024-06-02Star-
X-VILA: Cross-Modality Alignment for Large Language Model (by NVIDIA)arXiv2024-05-29-Wechat
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal ModelsarXiv2024-05-24Star-
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language ModelsarXiv2024-05-24--
LOVA3: Learning to Visual Question Answering, Asking and AssessmentarXiv2024-05-23Star-
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityarXiv2024-05-23StarProject Page
Chameleon: Mixed-Modal Early-Fusion Foundation Models (Meta)arXiv2024-05-16StarBlog
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-ExpertsarXiv2024-05-09StarProject Page
Dataset
Wechat
ImageInWords: Unlocking Hyper-Detailed Image Descriptions (Google)arXiv2024-05-05StarProject Page
Dataset
What matters when building vision-language models? (Idefics2)arXiv2024-05-03-
Collections
MANTIS: Interleaved Multi-Image Instruction TuningarXiv2024-05-02StarProject Page
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMsCVPR 2024 Workshop2024-04-23-Wechat
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language ModelsarXiv2024-04-19StarProject Page
Dataset
MoVA: Adapting Mixture of Vision Experts to Multimodal ContextarXiv2024-04-19Star-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language ModelsarXiv2024-04-18-Project Page
Project Page
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation? (LaDiC)NAACL 20242024-04-16Star-
AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception (AesExpert, AesMMIT Dataset)arXiv2024-04-15Star-
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models (Ferret-v2)arXiv2024-04-11--
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies (MiniCPM series)arXiv2024-04-09Star
Star
Blog
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs (Ferret-UI)arXiv2024-04-08--
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video UnderstandingCVPR 20242024-04-08StarProject Page
Koala: Key frame-conditioned long video-LLMCVPR 20242024-04-05StarProject Page
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual TokensarXiv2024-04-04StarProject Page
LongVLM: Efficient Long Video Understanding via Large Language ModelsarXiv2024-04-04Star-
InternVideo2: Scaling Foundation Models for Multimodal Video UnderstandingECCV 20242024-03-22Star-
VideoAgent: Long-form Video Understanding with Large Language Model as Agent (key frame)arXiv2024-03-15--
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training (Apple)arXiv2024-03-14--
UniCode: Learning a Unified Codebook for Multimodal Large Language ModelsarXiv2024-03-14--
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextarXiv2024-03-08-Project Page
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language ModelsarXiv2023-03-05Star-
RegionGPT: Towards Region Understanding Vision Language ModelCVPR 20242024-03-04-Project Page
All in an Aggregated Image for In-Image LearningarXiv2024-02-28Star-
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersCVPR 20242024-02-27StarProject Page
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different LanguagesarXiv2024-02-25--
LLMBind: A Unified Modality-Task Integration FrameworkarXiv2024-02-22--
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model (ALLaVA)arXiv2024-02-18StarDemo Page
Dataset
MobileVLM V2: Faster and Stronger Baseline for Vision Language ModelarXiv2024-02-06Star-
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile DevicesarXiv2023-12-28Star-
Gemini: A Family of Highly Capable Multimodal ModelsarXiv2023-12-19-Project Page
Osprey: Pixel Understanding with Visual Instruction TuningCVPR 20242023-12-15Star-
VILA: On Pre-training for Visual Language Models (NVIDIA, MIT)CVPR 20242023-12-12Star-
Vary: Scaling up the Vision Vocabulary for Large Vision-Language ModelsarXiv2023-12-11StarProject Page
Prompt Highlighter: Interactive Control for Multi-Modal LLMsCVPR 20242023-12-07StarProject Page
PixelLM: Pixel Reasoning with Large Multimodal ModelCVPR 20242023-12-04StarProject Page
APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsEMNLP 20232023-12-04StarProject Page
LLaMA-VID: An Image is Worth 2 Tokens in Large Language ModelsarXiv2023-11-28StarProject Page
Dataset
PG-Video-LLaVA: Pixel Grounding Large Video-Language ModelsarXiv2023-11-22Star-
ShareGPT4V: Improving Large Multi-Modal Models with Better CaptionsarXiv2023-11-21StarProject Page
LION : Empowering Multimodal Large Language Model with Dual-Level Visual KnowledgeCVPR 20242023-11-20StarProject Page
Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionarXiv2023-11-16Star-
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality CollaborationarXiv2023-11-07Star-
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learningarXiv2023-10-14StarProject Page
Ferret: Refer and Ground Anything Anywhere at Any Granularity (Ferret)ICLR 20242023-10-11Star-
Improved Baselines with Visual Instruction Tuning (LLaVA-1.5)arXiv2023-10-05StarProject Page
Aligning Large Multimodal Models with Factually Augmented RLHF (LLaVA-RLHF, MMHal-Bench (hallucination))arXiv2023-09-25StarProject Page
Dataset
MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningICLR 20242023-09-14Star-
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and BeyondarXiv2023-08-24StarProject Page
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages (VisCPM-Chat/Paint)ICLR 20242023-08-23Star-
SVIT: Scaling up Visual Instruction TuningarXiv2023-07-09StarDataset
Kosmos-2: Grounding Multimodal Large Language Models to the World (Kosmos-2, GrIT Dataset)arXiv2023-06-26StarDemo
Dataset
M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction TuningarXiv2023-06-07-Project Page
Dataset
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningNeurIPS 20232023-05-11Star-
MultiModal-GPT: A Vision and Language Model for Dialogue with HumansarXiv2023-05-08Star-
VPGTrans: Transfer Visual Prompt Generator across LLMsNeurIPS 20232023-05-02StarProject Page
mPLUG-Owl: Modularization Empowers Large Language Models with MultimodalityarXiv2023-04-27Star-
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsICLR 20242023-04-20StarProject Page
Visual Instruction Tuning (LLaVA)NeurIPS 20232023-04-17StarProject Page
Dataset
Language Is Not All You Need: Aligning Perception with Language Models (Kosmos-1)NeurIPS 20232023-02-27Star-
Multimodal Chain-of-Thought Reasoning in Language ModelsarXiv2023-02-02Star-
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsICML 20232023-01-30Star-
Flamingo: a Visual Language Model for Few-Shot LearningNeurIPS 20222022-04-29Star-

Omni Understanding

TitleVenueDateCodeSupplement
Ming-Omni: A Unified Multimodal Model for Perception and Generation (Ant Group)arXiv2025-06-11StarProject Page
Qwen2.5-Omni Technical ReportarXiv2025-03-26StarProject Page
Collections
PAVE: Patching and Adapting Video Large Language ModelsCVPR 20252025-03-25StarDataset
Baichuan-Omni-1.5 Technical ReportarXiv2025-01-26Star-
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement LearningarXiv2025-03-07StarModel
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferencearXiv2025-02-25StarProject Page
Collections
Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment (THU, Tencent Hunyuan, NTU S-Lab)arXiv2025-02-06StarProject Page
[Benchmark] WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs (Xiaohongshu, SJTU)arXiv2025-02-06StarProject Page
Dataset
Align Anything: Training All-Modality Models to Follow Instructions with Language FeedbackarXiv2024-12-20StarProject Page
Dataset
[Survey] From Specific-MLLM to Omni-MLLM: A Survey about the MLLMs alligned with Multi-Modality (HIT, Peng Cheng Lab)arXiv2024-12-16--
OMCAT: Omni Context Aware Transformer (OCTAV, OMCAT) (NVIDIA)arXiv2024-10-15-Project Page
Baichuan-Omni Technical ReportarXiv2024-10-11Star-
[Benchmark] OmniBench: Towards The Future of Universal Omni-Language ModelsarXiv2024-09-23StarProject Page
Dataset
OmniBind: Large-scale Omni Multimodal Representation via Binding SpacesarXiv2024-07-16StarProject Page
Explore the Limits of Omni-modal Pretraining at Scale (MiCo)arXiv2024-06-13StarProject Page
ViT-Lens: Towards Omni-modal Representations (TencentARC)CVPR 20242023-08-20StarProject Page
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetNeurIPS 20232023-05-29Star-
ImageBind: One Embedding Space To Bind Them AllCVPR 20232023-05-09StarProject Page
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and DatasetTPAMI 20242023-04-17StarProject Page

Unified Understanding and Generation

TitleVenueDateCodeSupplement
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal VelocitiesarXiv2025-05-26-Project Page
Wechat
MMaDA: Multimodal Large Diffusion Language Models (ByteDance Seed)arXiv2025-05-21StarDemo
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation (Apple)arXiv2025-05-20--
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation (ByteDance Seed)arXiv2025-05-08--
[Survey] Unified Multimodal Understanding and Generation Models: Advances, Challenges, and OpportunitiesarXiv2025-05-05Star-
Unified Reward Model for Multimodal Understanding and Generation (UnifiedReward) (Fudan, Shanghai AI Lab)arXiv2025-03-07StarProject Page
Dataset
UniTok: A Unified Tokenizer for Visual Generation and Understanding (ByteDance)arXiv2025-02-27StarProject Page
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling (by deepseek)arXiv2025-01-29Star-
LlamaFusion: Adapting Pretrained Language Models for Multimodal Generation (Meta)arXiv2024-12-19--
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingarXiv2024-12-12coming soon-
ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhancearXiv2024-12-09-Wechat
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation (TencentARC)arXiv2024-12-05Star-
Liquid: Language Models are Scalable Multi-modal Generators (Bytedance)arXiv2024-12-05StararXiv
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation (ByteDance)arXiv2024-12-04StarProject Page
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific HeadsICML 20252024-11-28Star-
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation (by deepseek)-2024-10-17Star-
Emu3: Next-Token Prediction is All You NeedarXiv2024-09-27StarProject Page
Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationarXiv2024-08-22StarProject Page
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal ModelICLR 2025 Oral2024-08-20Star-
An Image is Worth 32 Tokens for Reconstruction and Generation (TiTok, by ByteDance)arXiv2024-06-11StarProject Page
Mini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsarXiv2024-05-27StarProject Page
Collections
VITRON: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing-2024-04-25StarProject Page
YouTube
Wechat
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and GenerationarXiv2024-04-22Star-
AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingarXiv2024-02-19StarProject Page
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationarXiv2024-02-05StarProject Page
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and ActionarXiv2023-12-28StarProject Page
Generative Multimodal Models are In-Context Learners (Emu2)CVPR 20242023-12-20StarProject Page
CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any GenerationarXiv2023-11-30StarProject Page
LLMGA: Multimodal Large Language Model based Generation AssistantarXiv2023-11-27StarProject Page
VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and GenerationarXiv2023-12-14Star-
Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsICLR 20242023-10-04StarProject Page
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative VokensarXiv2023-10-03StarProject Page
DreamLLM: Synergistic Multimodal Comprehension and CreationICLR 20242023-09-20StarProject Page
NExT-GPT: Any-to-Any Multimodal LLMarXiv2023-09-11StarProject Page
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization (LaVIT)ICLR 20242023-09-09Star-
Planting a SEED of Vision in Large Language ModelICLR 20242023-07-16StarProject Page
Generative Pretraining in Multimodality (Emu1)ICLR 20242023-07-11Star-
Generating Images with Multimodal Language Models (GILL)NeurIPS 20232023-05-26StarProject Page
Any-to-Any Generation via Composable Diffusion (CoDi-1)NeurIPS 20232023-05-19StarProject Page
Grounding Language Models to Images for Multimodal Inputs and Outputs (FROMAGe)ICML 20232023-01-31StarProject Page

Coding / GUI

Vibe Coding / GUI / UI Design

TitleVenueDateCodeSupplement
[Model, Benchmark] UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation (THU, Zhipu)arXiv2025-11-14StarProject Page
[Agent] Computer-Use Agents as Judges for Generative User Interface (AUI) (Oxon, NUS Show Lab, Microsoft)arXiv2025-11-19StarProject Page
[Agent, Benchmark] Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development (TDDev)arXiv2025-09-29Star-
[Model] UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement LearningarXiv2025-09-02StarProject Page
Showcase
[Agent, Benchmark] FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI AgentsarXiv2025-08-12--
[Agent] ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal AgentsarXiv2025-07-30StarDemo
[Benchmark] ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation EvaluationarXiv2025-07-07StarProject Page
Dataset
[Benchmark] DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code GenerationarXiv2025-06-06StarDataset
[Model, Benchmark] WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from ScratcharXiv2025-05-06StarProject Page
Dataset
Dataset
[Benchmark] GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent (GTArena)arXiv2024-12-24StarDataset
[Benchmark] Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive PrototypingarXiv2024-11-05StarProject Page
Dataset
[Agent, Benchmark] Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design PrototypingNAACL 20252024-10-21StarProject Page
Dataset
[Dataset] WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsWWW 20252024-04-09StarProject Page
Dataset
[Benchmark] Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End EngineeringNAACL 20252024-03-05StarProject Page
Dataset

Diffusion MLLM

Multimodal Embedding/Retrieval

TitleVenueDateCodeSupplement
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought ReasoningEMNLP 20252025-09-25StarProject Page
Dataset
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual RetrievalarXiv2025-06-23-Project Page
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval (UNITE)arXiv2025-05-26StarProject Page
Collections
Wechat
[Benchmark] MIEB: Massive Image Embedding BenchmarkarXiv2025-04-14StarDemo
[Data, Model, Benchmark] IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal RetrievalarXiv2025-04-01Star-
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive LearningarXiv2025-03-04StarCollections
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval (CoT)arXiv2025-02-28--
[Model, Dataset] Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval (FDCA, FineCVR-1M)ICLR 20252025-02-26StarProject Page
[Benchmark, Model] MomentSeeker: A Comprehensive Benchmark and A Strong Baseline For Moment Retrieval Within Long Videos (MomentSeeker, V-Embedder) (Gaoling)arXiv2025-02-18StarDataset
[Data, Model, Benchmark] Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval (Vis-IR task, VIRA, UniSE, MVRB)arXiv2025-02-17--
[Data, Model] Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval (FineCVR-1M, FDCA)ICLR 20252025-01-23StarProject Page
[Benchmark] CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval (CaReBench)CVPR 20252024-12-31StarProject Page
Dataset
MINIMA: Modality Invariant Image MatchingCVPR 20252024-12-27StarDemo
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs (Tongyi Lab)CVPR 20252024-12-22StarCollections
✨ [Dataset, Model] MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval (MagaPairs, BGE-VL)arXiv2024-12-19StarDataset
Wechat
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval (OSrCIR) (CoT)CVPR 20252024-12-15Star-
LamRA: Large Multimodal Model as Your Advanced Retrieval AssistantarXiv2024-12-02StarProject Page
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs (NVIDIA)ICLR 2025 Poster2024-11-04-Model
OMCAT: Omni Context Aware TransformerarXiv2024-10-15-Project Page
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks (TIGER Lab)arXiv2024-10-07StarProject Page
Dataset
Dataset
Collections
E5-V: Universal Embeddings with Multimodal Large Language ModelsarXiv2024-07-17Star-
OmniBind: Large-scale Omni Multimodal Representation via Binding SpacesarXiv2024-07-16StarProject Page
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models (NVIDIA)ICLR 20252024-05-27-Collections
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions (Google DeepMind)ICML 2024 Oral2024-03-28StarProject Page
DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation ModelsNAACL2024-04-07--
Composed Video Retrieval via Enriched Context and Discriminative EmbeddingsCVPR 20242024-03-25Star-
UniIR: Training and Benchmarking Universal Multimodal Information Retrievers (TIGER Lab)ECCV 20242023-11-28StarProject Page
Dataset
CoVR-2: Automatic Data Construction for Composed Video Retrieval&CoVR: Learning Composed Video Retrieval from Web Video CaptionsTPAMI 2024 & AAAI 20242023-08-23StarProject Page
Dataset
Dataset
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetNeurIPS 20232023-05-29Star-
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and DatasetTPAMI20242023-04-17StarProject Page
✨ (QB-Norm) Cross Modal Retrieval with Querybank NormalisationCVPR 20222021-12-23StarProject Page
✨ (DSL) Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax LossarXiv2021-09-09Star-

Image Understanding Benchmark

Video Understanding Benchmark

TitleVenueDateCodeSupplement
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsACL 20262026-03StarProject Page
Dataset
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic VideosACL 2025 Main2025-05-26StarProject Page
Dataset
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation ModelsarXiv2024-10-30StarDataset
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark (AuroraCap, VDC)arXiv2024-10-24StarProject Page
Dataset
Dataset
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video ModelsarXiv2024-10-14StarProject Page
Dataset
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingNeurIPS 20242024-09-26StarProject Page
Dataset
Dataset
Tarsier: Recipes for Training and Evaluating Large Video Description Models (Tarsier, Dream1k) (ByteDance)arXiv2024-07-30StarDataset
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning (VideoVista)arXiv2024-06-17StarProject Page
Dataset
VELOCITI: Can Video-Language Models Bind Semantic Concepts through Time?arXiv2024-06-16StarProject Page
Dataset
MLVU: A Comprehensive Benchmark for Multi-Task Long Video UnderstandingarXiv2024-06-06StarDataset
Dataset
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis (Video-MME)arXiv2024-05-31StarProject Page
Dataset
TempCompass: Do Video LLMs Really Understand Videos?arXiv2024-03-01StarProject Page
Dataset
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark (MVBench, VideoChat2)CVPR 2024 Highlight2023-11-28StarDataset
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language UnderstandingNeurIPS 20232023-08-17StarProject Page
Dataset
Perception Test: A Diagnostic Benchmark for Multimodal Video Models (Perception Test, by Google DeepMind)NeurIPS 20232023-05-23StarProject Page
Dataset

Audio

Multimodal Dialogue

Multimodal Learning

Image Generation

TitleVenueDateCodeSupplement
ImageBench V1: Text-to-Image Benchmark with VLM Judges (live benchmark; 40+ models, 192 prompts, all outputs published)WebsiteLive-Project Page Methodology
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and EditingarXiv2025-03-13Star-
OmniGen: Unified Image GenerationarXiv2024-09-17StarProject Page
Demo
Wechat
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens (by Kaiming He, DeepMind, MIT)arXiv2024-10-17--
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (Lumina-T2X, Flag-DiT) (Text2Any)arXiv2024-05-09StarYouTube
Wechat
FreeU: Free Lunch in Diffusion U-Net (FreeU, by Ziwei Liu)CVPR 2024 Oral2023-09-20StarProject Page
YouTube
Demo
Lazy Diffusion Transformer for Interactive Image EditingarXiv2024-04-18-Project Page
Salient Object-Aware Background Generation using Text-Guided Diffusion ModelsCVPR 2024 Workshop2024-04-15Star-
HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingarXiv2024-04-15StarProject Page
Dataset
Demo
UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark (UNIAA-LLaVA, UNIAA-Bench)arXiv2024-04-15--
PMG: Personalized Multimodal Generation with Large Language ModelsWWW 20242024-04-07--
Identity Decoupling for Multi-Subject Personalization of Text-to-Image ModelsarXiv2024-04-05StarProject Page
Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsCVPR 20242024-04-05--
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (VAR)arXiv2024-04-03StarProject Page
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation (HuaWei, Enze Xie)arXiv2024-03-07StarProject Page
Multi-LoRA Composition for Image GenerationarXiv2024-02-26StarProject Page
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models (HuaWei, Enze Xie)arXiv2024-01-10StarProject Page
Brush Your Text: Synthesize Any Scene Text on Images via Diffusion ModelAAAI 20242023-12-19Star-
SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models (Tencent Xintao Wang)arXiv2023-12-11StarProject Page
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction FollowingarXiv2023-12-11StarProject Page
Emu Edit: Precise Image Editing via Recognition and Generation TasksarXiv2023-11-16-Project Page
BeautifulPrompt: Towards Automatic Prompt Engineering for Text-to-Image SynthesisEMNLP 20232023-11-12Starzhihu
AnyText: Multilingual Visual Text Generation And EditingICLR 20242023-11-06Star-
EasyGen: Easing Multimodal Generation with a Bidirectional Conditional Diffusion Model and LLMsarXiv2023-10-13Star-
Mini-DALLE3: Interactive Text to Image by Prompting Large Language ModelsarXiv2023-10-11StarProject Page
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis (HuaWei, Enze Xie)ICLR 2024 Spotlight2023-09-30StarProject Page
Dataset
Usage Diffusers
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion ModelsarXiv2023-08-13StarProject Page
Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsarXiv2023-10-04StarProject Page
Improving Image Generation with Better Captions (DALL-E 3)OpenAI2023--
Scaling up GANs for Text-to-Image Synthesis (GigaGAN)CVPR 20232023-05-09StarProject Page
Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)ICCV 20232023-02-10Star-
Scalable Diffusion Models with Transformers (DiT)ICCV 20232022-12-19StarProject Page
InstructPix2Pix: Learning to Follow Image Editing InstructionsCVPR 20232022-11-17StarProject Page
All are Worth Words: A ViT Backbone for Diffusion Models (U-ViT, first Diffsuion Transformer) (RUC, Chongxuan Li)CVPR 20232022-09-25Star-
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven GenerationCVPR 20232022-08-25StarProject Page
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding (Imagen)NeurIPS 20222022-05-23StarProject Page
Hierarchical Text-Conditional Image Generation with CLIP Latents (DALL-E 2)OpenAI2022-04-13Star-
High-Resolution Image Synthesis with Latent Diffusion Models (LDM, Stable Diffusion)CVPR 20222021-12-20Star-
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsICML 20222021-12-20Star-
NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtionECCV 20222021-11-24Star-
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsICLR 20222021-08-02StarProject Page
CogView: Mastering Text-to-Image Generation via TransformersNeurIPS 20212021-05-26Star-
Zero-Shot Text-to-Image Generation (DALL-E 1)ICML 20212021-02-24StarProject Page
Taming Transformers for High-Resolution Image Synthesis (VQ-GAN)CVPR 20212020-12-17StarProject Page

Video Generation

TitleVenueDateCodeSupplement
ReCamMaster: Camera-Controlled Generative Rendering from A Single VideoICCV20252025-03-14StarProject Page
Dataset
[Dataset] Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video SpecialistsarXiv2025-02-10StarProject Page
Dataset
Demo
LiFT: Leveraging Human Feedback for Text-to-Video Model AlignmentarXiv2024-12-06StarProject Page
[Dataset] VidGen-1M: A Large-Scale Dataset for Text-to-video GenerationarXiv2024-08-05StarProject Page
Dataset
[Dataset] MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions (Mira)arXiv2024-07-08Video GenerationStar
VIMI: Grounding Video Generation through Multi-modal InstructionarXiv2024-07-08StarProject Page
[Dataset] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationICLR 20252024-07-02StarProject Page
Dataset
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (Lumina-T2X, Flag-DiT) (Text2Any)arXiv2024-05-09StarYouTube
Wechat
StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text (Long Video Generation)arXiv2024-03-21StarProject Page
YouTube
Demo
Wechat
AnyV2V: A Plug-and-Play Framework For Any Video-to-Video Editing TasksarXiv2024-03-21StarProject Page
Demo
Demo Page
Wechat
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation (FRESCO) (NTU, Ziwei Liu)CVPR 20242024-03-19StarProject Page
Latte: Latent Diffusion Transformer for Video Generation (Latte) (NTU, Ziwei Liu)arXiv2024-01-05StarProject Page
FreeInit: Bridging Initialization Gap in Video Diffusion Models (FreeInit) (NTU, Ziwei Liu)arXiv2023-12-12StarProject Page
YouTube
Demo
VideoBooth: Diffusion-based Video Generation with Image Prompts (VideoBooth) (NTU, Ziwei Liu)arXiv2023-12-01StarProject Page
VBench: Comprehensive Benchmark Suite for Video Generative Models [Benchmark] (VBench) (NTU, Ziwei Liu)CVPR 20242023-11-29StarProject Page
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (SVD)arXiv2023-11-25StarProject Page
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction (NTU, Ziwei Liu)ICLR 20242023-10-31StarProject Page
FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling (FreeNoise) (NTU, Ziwei Liu)ICLR 20242023-10-23StarProject Page
LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models (LaVie) (NTU, Ziwei Liu)arXiv2023-09-26StarProject Page

Multimodal Dataset

TitleVenueDateCodeSupplement
[Benchmark & Dataset] VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsarXiv2025-04-21Visual Reasoning Benchmark & Dataset🌐 Homepage
🤗 Benchmark
🤗 Train Data
💻 Train Code
[Dataset] Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video SpecialistsarXiv2025-02-10StarProject Page
Dataset
Demo
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction DataarXiv2024-10-24-Dataset
LVD-2M: A Long-take Video Dataset with Temporally Dense CaptionsNeurIPS 20242024-10-14StarProject Page
VidGen-1M: A Large-Scale Dataset for Text-to-video GenerationarXiv2024-08-05StarProject Page
Dataset
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions (Mira)arXiv2024-07-08Video GenerationStar
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationICLR 20252024-07-02StarProject Page
Dataset
GUIDE: A Guideline-Guided Dataset for Instructional Video ComprehensionIJCAI 20242024-06-26-Project Page
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion TokensarXiv2024-06-17StarCollections
Blog
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and GenerationarXiv2024-06-15Star-
What If We Recaption Billions of Web Images with LLaMA-3? (Recap-DataComp-1B)arXiv2024-06-12StarProject Page
Dataset
[Dataset
TextSquare: Scaling up Text-Centric Visual Instruction TuningarXiv2024-04-19Visual Instruction Tuning-
HQ-Edit: A High-Quality Dataset for Instruction-based Image EditingarXiv2024-04-15Instruction Image EditingStar
Project Page
Dataset
Demo
AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception (AesExpert, AesMMIT Dataset)arXiv2024-04-15Aesthetic Multi-Modality Instruction TuningStar
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality TeachersCVPR 20242024-02-29video-captionStar
Project Page
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language ModelarXiv2024-02-18GPT4V-synthesized DataStar
Demo Page
Dataset
STICKERCONV: Generating Multimodal Empathetic Responses from ScratchACL 2024 Main2024-01-20Multimodal Empathetic DialogueStar
Project Page
Dataset
Wechat
zhihu
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationICLR 20242023-07-13StarDataset
Dataset
SVIT: Scaling up Visual Instruction TuningarXiv2023-07-09Instruction TuningStar
Dataset
Kosmos-2: Grounding Multimodal Large Language Models to the World (Kosmos-2, GrIT Dataset)arXiv2023-06-26Grounded image-text pairsStar
Demo
Dataset
M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction TuningarXiv2023-06-07Instruction TuningProject Page
Dataset
Visual Instruction Tuning (LLaVA)NeurIPS 20232023-04-17Instruction TuningStar
Project Page
Dataset
Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with TextNeurIPS D&B 20232023-04-14Interleaved Image-TextStar
AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content Creation (Wiki)TREC 2023 Workshop2023-04-04StarDataset
Dataset
TikTalk: A Multi-Modal Dialogue Dataset for Real-World ChitchatACM MM 20232023-01-14Multimodal DialogueStar
MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationACL 20232022-11-10Multimodal DialogueStar
LAION-5B: An open large-scale dataset for training next generation image-text modelsNeurIPS 20222022-10-16Image-Text PairsProject Page
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text PairsNeurIPS Workshop 20212021-11-03Image-Text PairsProject Page
MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsACM SIGIR 20212021-07Multimodal DialogueStar
PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text ModelingACL 20212021-07-06Open-domain Multimodal DialogueStar
Image-Chat: Engaging Grounded ConversationsACL 20202018-11-02Multimodal DialogueProject Page

Multimodal Survey

TitleVenueDateSupplementLatest Update
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and ApplicationsTMLR 20262026-05--
Discrete Diffusion in Large Language and Multimodal Models: A SurveyarXiv2025-06-16Star-
A Survey on Bridging VLMs and Synthetic DataOpenReview2025-05-16Star
Github Page
-
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and OpportunitiesarXiv2025-05-05Star-
From Specific-MLLM to Omni-MLLM: A Survey about the MLLMs alligned with Multi-Modality (HIT, Peng Cheng Lab)arXiv2024-12-16-
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding (TikTok)arXiv2024-09-27-
Video Diffusion Models: A SurveyarXiv2024-05-06Star-
Theoretical research on generative diffusion models: an overviewarXiv2024-04-13--
A Review of Multi-Modal Large Language and Vision ModelsarXiv2024-03-28--
The (R)Evolution of Multimodal Large Language Models: A SurveyarXiv2024-02-19--
MM-LLMs: Recent Advances in MultiModal Large Language ModelsarXiv2024-01-24-2024-02-20
Multimodal Large Language Models: A SurveyIEEE BigData 20232023-11-22--
Multimodal Foundation Models: From Specialists to General-Purpose AssistantsCVPR 20232023-09-18--
Understanding Deep Learning-2023--
Large Multimodal Models: Notes on CVPR 2023 TutorialCVPR 20232023-06-26--
A Survey on Multimodal Large Language ModelsarXiv2023-06-23-2024-04-01
Multimodal Deep LearningarXiv2023-01-12--
Diffusion Models: A Comprehensive Survey of Methods and ApplicationsACM Computing Surveys2022-09-02-2024-02-06
Multimodal Learning with Transformers: A SurveyIEEE TPAMI 20232022-01-13-2023-05-10
Multimodal Machine Learning: A Survey and TaxonomyIEEE PAMI 20192017-05-26-2017-08-01
deep-learning
large-multimodal-models
multimodal
multimodal-data
multimodal-deep-learning
multimodal-dialogue
multimodal-large-language-models
multimodal-learning

Contributors

friedrichor

94 commits

mghiasvand1

3 commits

dh7

1 commits

xwy-bit

1 commits