Yuan-ManX/ai-multimodal-timeline

Here we will track the latest AI Multimodal Models, including Multimodal Foundation Models, LLM, Agent, Audio, Image, Video, Music and 3D content. 🔥

36

165 commits

updated Feb 4, 2025

See the code

README

AI Multimodal Timeline

AI Multimodal Timeline

Here we will track the latest AI Multimodal Models, including Multimodal Foundation Model, LLM, Agent, Audio, Image, Video, Music and 3D content. 🔥

Table of Contents

Project List

Multimodal Model

DateSourceDescriptionPaperModel
2025-01MILSMILS: LLMs can see and hear without any training.arXiv
2024-11OasisOasis is an interactive world model developed by Decart and Etched. Based on diffusion transformers, Oasis takes in user keyboard input and generates gameplay in an autoregressive manner.Hugging Face
2024-10UnboundedUnbounded: A Generative Infinite Game of Character Life Simulation.arXivWebsite
2024-10JanusJanus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.arXivHugging Face
2024-09LLaVA-3DLLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.arXiv
2024-09Emu3Emu3: Next-Token Prediction is All You Need.Hugging Face
2024-09MoshiMoshi: a speech-text foundation model for real time dialogue.Hugging Face
2024-09Qwen2-VLQwen2-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.Hugging Face
2024-08EagleEagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.arXiv
2024-08Mini-OmniMini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.arXivHugging Face
2024-08GameNGenGameNGen - Diffusion Models Are Real-Time Game Engines.arXiv
2024-08SapiensSapiens: Foundation for Human Vision Models.arXiv
2024-08Show-oShow-o: One Single Transformer to Unify Multimodal Understanding and Generation.arXiv
2024-08LLaVA-OneVisionLLaVA-OneVision: Easy Visual Task Transfer.arXivHugging Face
2024-08AI ScientistThe AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.arXiv
2024-08Mini-MonkeyMini-Monkey: Multi-Scale Adaptive Cropping for Multimodal Large Language Models.arXiv
2024-08VITAVITA: Towards Open-Source Interactive Omni Multimodal LLM.arXiv
2024-08Lumina-mGPTLumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.arXiv
2024-07Any2PointAny2Point: Empowering Any-modality Large Models for Efficient 3D Understanding.arXiv
2024-07SOLOSOLO: A Single Transformer for Scalable Vision-Language Modeling.arXiv
2024-07KangarooKangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.Hugging Face
2024-07SEED-StorySEED-Story: Multimodal Long Story Generation with Large Language Model.arXivHugging Face
2024-07VTA-LDMVideo-to-Audio Generation with Hidden Alignment.arXivHugging Face
2024-07Qwen2-AudioQwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.arXiv
2024-07MoshiMoshi is an experimental conversational AI.Website
2024-07AnoleAnole: An Open, Autoregressive and Native Multimodal Models for Interleaved Image-Text Generation.Hugging Face
2024-06Cambrian-1A Fully Open, Vision-Centric Exploration of Multimodal LLMs.arXivHugging Face
2024-06EVF-SAMEVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.arXivHugging Face
2024-06MINT-1TScaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens.arXiv
2024-06OmniTokenizerA Joint Image-Video Tokenizer for Visual Generation.arXivWebsite
2024-06ml-4mA framework for training any-to-any multimodal foundation models.arXivWebsite
2024-06LongVALong Context Transfer from Language to Vision.arXivHugging Face
2024-06VideoLLaMA 2Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.arXivHugging Face
2024-05ManyICLMany-Shot In-Context Learning in Multimodal Foundation Models.arXiv
2024-05Contrastive ALignment (CAL)Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment.arXiv
2024-05GromaGrounded Multimodal Large Language Model with Localized Visual Tokenization.arXivHugging Face
2024-05CogVLM2GPT4V-level open-source multi-modal model based on Llama3-8B.Hugging Face
2024-05ChameleonMixed-Modal Early-Fusion Foundation Models.arXiv
2024-05Lumina-T2XTransforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.arXivHugging Face
2024-05MiniCPM-Llama3-V 2.5MiniCPM-Llama3-V 2.5 is the latest model in the MiniCPM-V series. The model is built on SigLip-400M and Llama3-8B-Instruct with a total of 8B parameters.Hugging Face
2024-05GeminiBuild with state-of-the-art generative models and tools to make AI helpful for everyone.API
2024-05GPT-4oGPT-4o (“o” for “omni”) is a step towards much more natural human-computer interaction—it accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs.API
2024-04MyGODiscrete Modality Information as Fine-Grained Tokens for Multi-modal Knowledge Graph Completion.arXiv
2024-04InternLM-XComposer2InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension.arXivHugging Face
2024-02AnyGPTAnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.arXiv
2024-01MMVPExploring the Visual Shortcomings of Multimodal LLMs.arXiv
2023-12V*Guided Visual Search as a Core Mechanism in Multimodal LLMs.arXiv
2023-12Tokenize AnythingTokenize Anything via Prompting.arXivHugging Face
2023-12VILAVILA: On Pre-training for Visual Language Models.arXivHugging Face
2023-11LEOAn Embodied Generalist Agent in 3D World.arXivWebsite
2023-11ShareGPT4VImproving Large Multi-Modal Models with Better Captions.arXivHugging Face
2023-11Video-LLaVALearning United Visual Representation by Alignment Before Projection.arXivHugging Face
2023-10LanguageBindExtending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.arXivHugging Face
2023-07EmuEmu: Generative Multimodal Models from BAAI.arXivHugging Face
2023-05ImageBindOne Embedding Space To Bind Them All.arXivWebsite
2022-11EVAEVA: Visual Representation Fantasies from BAAI.arXivHugging Face

^ Back to Contents ^

LLM

DateSourceDescriptionPaperModel
2024-08LongWriterLongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.arXivHugging Face
2024-07DCLMDataComp for Language ModelsarXivHugging Face
2024-07Index-1.9BA SOTA lightweight multilingual LLMHugging Face
2024-06Claude 3.5 SonnetClaude 3.5 SonnetAPI
2024-06Nemotron-4Nemotron-4-340B-Instruct is a large language model (LLM) that can be used as part of a synthetic data generation pipeline to create training data that helps researchers and developers build their own LLMs.arXivHugging Face
2024-06Qwen2Qwen2 is the large language model series developed by Qwen team, Alibaba Cloud.Hugging Face
2024-04Llama 3Meta Llama 3 is the next generation of our state-of-the-art open source large language model.Hugging Face
2024-03Claude 3Talk with Claude, an AI assistant from Anthropic.API
2024-03Grok-1The weights and architecture of our 314 billion parameter Mixture-of-Experts model, Grok-1.Hugging Face
2023-11MixtralOpen and portable generative AI for devs and businesses.arXivHugging Face
2023-09Baichuan 2A series of large language models developed by Baichuan Intelligent Technology.Hugging Face
2023-07GPT-4GPT-4 is OpenAI’s most advanced system, producing safer and more useful responses.API

^ Back to Contents ^

Agent

DateSourceDescriptionPaperModel
2024-10TEN AgentTEN Agent is the world’s first real-time multimodal agent integrated with the OpenAI Realtime API, RTC, and features weather checks, web search, vision, and RAG capabilities.Website
2024-08TwitterTwitter Personality is a web application that analyzes your Twitter handle to create a personalized personality profile using Wordware AI Agent.Website
2024-08MindSearch🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT).
2024-08MMRoleMMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents.arXiv
2024-08Agent KAn autoagentic AGI that is self-evolving and modular.
2024-08LangGraph StudioLangGraph Studio offers a new way to develop LLM applications by providing a specialized agent IDE that enables visualization, interaction, and debugging of complex agentic applications.
2024-07LLama Agentic SystemAgentic components of the Llama Stack APIs.
2024-07TaskGenA Task-based agentic framework building on StrictJSON outputs by LLM agents.
2024-07IoAAn open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity.
2024-07OmAgentA multimodal agent framework for solving complex tasks.arXiv
2024-06GraphRAGA modular graph-based Retrieval-Augmented Generation (RAG) system.Website
2024-06Mixture of Agents (MoA)Mixture-of-Agents Enhances Large Language Model Capabilities.arXiv
2024-06Buffer of ThoughtsThought-Augmented Reasoning with Large Language Models.arXiv
2024-06Translation AgentAgentic translation using reflection workflow.
2024-06Atomic AgentsThe Atomic Agents framework is designed to be modular, extensible, and easy to use.
2024-05PipecatOpen Source framework for voice and multimodal conversational AI.
2024-02V-IRLGrounding Virtual Intelligence in Real Life.arXiv

^ Back to Contents ^

Audio

Audio/Text-to-Speech

DateSourceDescriptionPaperModel
2024-07CosyVoiceMulti-lingual large voice generation model, providing inference, training and deployment full-stack ability.
2024-06DEX-TTSDiffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability.arXivWebsite
2024-05ChatTTSChatTTS is a text-to-speech model designed specifically for dialogue scenario such as LLM assistant.
2023-06StyleTTS 2Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.arXivHugging Face

Audio/Automatic Speech Recognition

DateSourceDescriptionPaperModel
2024-07SenseVoiceSenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED).Hugging Face
2024-05TeleSpeech-ASRLarge speech model-super multi-dialect ASR.Hugging Face
2022-12WhisperWhisper is a general-purpose speech recognition model.arXivAPI

Audio/Audio Generation

DateSourceDescriptionPaperModel
2024-07FoleyCrafterFoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.arXivHugging Face
2024-06SEE-2-SOUNDZero-Shot Spatial Environment-to-Spatial Sound.arXiv
2024-05Make-An-Audio 3Transforming Text into Audio via Flow-based Large Diffusion Transformers.arXivHugging Face

^ Back to Contents ^

Image

DateSourceDescriptionPaperModel
2024-09StoryMakerStoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation.arXivHugging Face
2024-08CSGOCSGO: Content-Style Composition in Text-to-Image Generation.arXiv
2024-08FLUXThis repo contains minimal inference code to run text-to-image and image-to-image with our Flux latent rectified flow transformers.Hugging Face
2024-08Segment Anything Model 2 (SAM 2)SAM 2: Segment Anything in Images and Videos.arXivHugging Face
2024-07CatVTONCatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models.arXivHugging Face
2024-07UltraEditUltraEdit: Instruction-based Fine-Grained Image Editing at Scale.arXivHugging Face
2024-07UltraPixelUltraPixel: Advancing Ultra-High-Resolution Image Synthesis to New Peaks.arXiv
2024-07PaintsUndoPaintsUndo: A Base Model of Drawing Behaviors in Digital Paintings.
2024-07KolorsKolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis.Hugging Face
2024-06Depth Anything V2Depth Anything V2.arXivHugging Face
2024-06AutoStudioCrafting Consistent Subjects in Multi-turn Interactive Image Generation.arXiv
2024-06MimicBrushZero-shot Image Editing with Reference Imitation.arXivHugging Face
2024-06LlamaGenAutoregressive Model Beats Diffusion: Llama for Scalable Image Generation.arXivHugging Face
2024-05OmostOmost is a project to convert LLM's coding capability to image generation (or more accurately, image composing) capability.Hugging Face
2024-05Hunyuan-DiTA Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.arXivHugging Face
2024-02MIGCMIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis.arXiv
2023-10DALL·E 3DALL·E is a AI system that can create realistic images and art from a description in natural language.API

^ Back to Contents ^

Video

DateSourceDescriptionPaperModel
2024-11LTX-VideoLTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time.Hugging Face
2024-09MIMOMIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling.arXivWebsite
2024-09DrawingSpinUpDrawingSpinUp: 3D Animation from Single Character Drawings.arXivWebsite
2024-09ViewCrafterViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis.arXivWebsite
2024-08CogVideoXCogVideoX is an open-source version of the video generation model, which is homologous to 清影.Hugging Face
2024-07ToraTora: Trajectory-oriented Diffusion Transformer for Video Generation.arXivWebsite
2024-06DiffutoonHigh-Resolution Editable Toon Shading via Diffusion Models.arXivWebsite
2024-05Video-MMEThe First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.
2024-05Video-of-ThoughtVideo-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition.Website
2024-05MOFA-VideoMOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model.arXivHugging Face
2024-05MotionLLMUnderstanding Human Behaviors from Human Motions and Videos.arXiv
2024-05ViduVidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.arXiv
2024-02SoraSora is an AI model that can create realistic and imaginative scenes from text instructions.Technical Report
2023-11PikaPika is the idea-to-video platform that sets your creativity in motion.
2023-03RunwayRunway is an applied AI research company shaping the next era of art, entertainment and human creativity.

^ Back to Contents ^

Music

DateSourceDescriptionPaperModel
2025-01YuEYuE: Open Full-song Generation Foundation Model, something similar to Suno.ai but open.Hugging Face
2024-05Diff-BGMA Diffusion Model for Video Background Music Generation.arXiv
2024-04UdioUdio - AI Music GeneratorWebsite
2023-12SunoSuno is building a future where anyone can make great music.Website
2023-12Soundry AIGenerative AI tools including text-to-sound and infinite sample packs.Website
2023-12SonautoSonauto is an AI music editor that turns prompts, lyrics, or melodies into full songs in any style.Website

^ Back to Contents ^

3D

DateSourceDescriptionPaperModel
2025-01Hunyuan3D 2.0Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation.arXivHugging Face
2024-11Hunyuan3D-1.0Hunyuan3D-1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation.arXivHugging Face
2024-093DTopia-XL3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion.arXivWebsite
2024-08ShapeSplatShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining.arXiv
2024-08SF3DSF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement.arXivHugging Face
2024-07HoloDreamerHoloDreamer: Holistic 3D Panoramic World Generation from Text Descriptions.arXivWebsite
2024-07DreamCatalystDreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation.arXivWebsite
2024-07CharacterGenCharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization.arXivWebsite
2024-07GALA3DGALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.arXivWebsite
2024-06Unique3DHigh-Quality and Efficient 3D Mesh Generation from a Single Image.arXivHugging Face
2024-06DreamGaussian4DGenerative 4D Gaussian Splatting.arXivHugging Face
2024-03GaussCtrlGaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing.arXiv
2024-03GaussianCubeA Structured and Explicit Radiance Representation for 3D Generative Modeling.arXivHugging Face
2024-03TripoSRFast 3D Object Reconstruction from a Single Image.arXivHugging Face

^ Back to Contents ^

ai
ai-agents
deeplearning-ai
llm
multi-modal
multimodal
multimodal-deep-learning

Contributors

Yuan-ManX

165 commits

Yuan-ManX/ai-multimodal-timeline

Here we will track the latest AI Multimodal Models, including Multimodal Foundation Models, LLM, Agent, Audio, Image, Video, Music and 3D content. 🔥

36

165 commits

updated Feb 4, 2025

See the code

README

AI Multimodal Timeline

AI Multimodal Timeline

Here we will track the latest AI Multimodal Models, including Multimodal Foundation Model, LLM, Agent, Audio, Image, Video, Music and 3D content. 🔥

Table of Contents

Project List

Multimodal Model

DateSourceDescriptionPaperModel
2025-01MILSMILS: LLMs can see and hear without any training.arXiv
2024-11OasisOasis is an interactive world model developed by Decart and Etched. Based on diffusion transformers, Oasis takes in user keyboard input and generates gameplay in an autoregressive manner.Hugging Face
2024-10UnboundedUnbounded: A Generative Infinite Game of Character Life Simulation.arXivWebsite
2024-10JanusJanus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.arXivHugging Face
2024-09LLaVA-3DLLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.arXiv
2024-09Emu3Emu3: Next-Token Prediction is All You Need.Hugging Face
2024-09MoshiMoshi: a speech-text foundation model for real time dialogue.Hugging Face
2024-09Qwen2-VLQwen2-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.Hugging Face
2024-08EagleEagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.arXiv
2024-08Mini-OmniMini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.arXivHugging Face
2024-08GameNGenGameNGen - Diffusion Models Are Real-Time Game Engines.arXiv
2024-08SapiensSapiens: Foundation for Human Vision Models.arXiv
2024-08Show-oShow-o: One Single Transformer to Unify Multimodal Understanding and Generation.arXiv
2024-08LLaVA-OneVisionLLaVA-OneVision: Easy Visual Task Transfer.arXivHugging Face
2024-08AI ScientistThe AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.arXiv
2024-08Mini-MonkeyMini-Monkey: Multi-Scale Adaptive Cropping for Multimodal Large Language Models.arXiv
2024-08VITAVITA: Towards Open-Source Interactive Omni Multimodal LLM.arXiv
2024-08Lumina-mGPTLumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.arXiv
2024-07Any2PointAny2Point: Empowering Any-modality Large Models for Efficient 3D Understanding.arXiv
2024-07SOLOSOLO: A Single Transformer for Scalable Vision-Language Modeling.arXiv
2024-07KangarooKangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.Hugging Face
2024-07SEED-StorySEED-Story: Multimodal Long Story Generation with Large Language Model.arXivHugging Face
2024-07VTA-LDMVideo-to-Audio Generation with Hidden Alignment.arXivHugging Face
2024-07Qwen2-AudioQwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.arXiv
2024-07MoshiMoshi is an experimental conversational AI.Website
2024-07AnoleAnole: An Open, Autoregressive and Native Multimodal Models for Interleaved Image-Text Generation.Hugging Face
2024-06Cambrian-1A Fully Open, Vision-Centric Exploration of Multimodal LLMs.arXivHugging Face
2024-06EVF-SAMEVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.arXivHugging Face
2024-06MINT-1TScaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens.arXiv
2024-06OmniTokenizerA Joint Image-Video Tokenizer for Visual Generation.arXivWebsite
2024-06ml-4mA framework for training any-to-any multimodal foundation models.arXivWebsite
2024-06LongVALong Context Transfer from Language to Vision.arXivHugging Face
2024-06VideoLLaMA 2Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.arXivHugging Face
2024-05ManyICLMany-Shot In-Context Learning in Multimodal Foundation Models.arXiv
2024-05Contrastive ALignment (CAL)Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment.arXiv
2024-05GromaGrounded Multimodal Large Language Model with Localized Visual Tokenization.arXivHugging Face
2024-05CogVLM2GPT4V-level open-source multi-modal model based on Llama3-8B.Hugging Face
2024-05ChameleonMixed-Modal Early-Fusion Foundation Models.arXiv
2024-05Lumina-T2XTransforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.arXivHugging Face
2024-05MiniCPM-Llama3-V 2.5MiniCPM-Llama3-V 2.5 is the latest model in the MiniCPM-V series. The model is built on SigLip-400M and Llama3-8B-Instruct with a total of 8B parameters.Hugging Face
2024-05GeminiBuild with state-of-the-art generative models and tools to make AI helpful for everyone.API
2024-05GPT-4oGPT-4o (“o” for “omni”) is a step towards much more natural human-computer interaction—it accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs.API
2024-04MyGODiscrete Modality Information as Fine-Grained Tokens for Multi-modal Knowledge Graph Completion.arXiv
2024-04InternLM-XComposer2InternLM-XComposer2 is a groundbreaking vision-language large model (VLLM) excelling in free-form text-image composition and comprehension.arXivHugging Face
2024-02AnyGPTAnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.arXiv
2024-01MMVPExploring the Visual Shortcomings of Multimodal LLMs.arXiv
2023-12V*Guided Visual Search as a Core Mechanism in Multimodal LLMs.arXiv
2023-12Tokenize AnythingTokenize Anything via Prompting.arXivHugging Face
2023-12VILAVILA: On Pre-training for Visual Language Models.arXivHugging Face
2023-11LEOAn Embodied Generalist Agent in 3D World.arXivWebsite
2023-11ShareGPT4VImproving Large Multi-Modal Models with Better Captions.arXivHugging Face
2023-11Video-LLaVALearning United Visual Representation by Alignment Before Projection.arXivHugging Face
2023-10LanguageBindExtending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.arXivHugging Face
2023-07EmuEmu: Generative Multimodal Models from BAAI.arXivHugging Face
2023-05ImageBindOne Embedding Space To Bind Them All.arXivWebsite
2022-11EVAEVA: Visual Representation Fantasies from BAAI.arXivHugging Face

^ Back to Contents ^

LLM

DateSourceDescriptionPaperModel
2024-08LongWriterLongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.arXivHugging Face
2024-07DCLMDataComp for Language ModelsarXivHugging Face
2024-07Index-1.9BA SOTA lightweight multilingual LLMHugging Face
2024-06Claude 3.5 SonnetClaude 3.5 SonnetAPI
2024-06Nemotron-4Nemotron-4-340B-Instruct is a large language model (LLM) that can be used as part of a synthetic data generation pipeline to create training data that helps researchers and developers build their own LLMs.arXivHugging Face
2024-06Qwen2Qwen2 is the large language model series developed by Qwen team, Alibaba Cloud.Hugging Face
2024-04Llama 3Meta Llama 3 is the next generation of our state-of-the-art open source large language model.Hugging Face
2024-03Claude 3Talk with Claude, an AI assistant from Anthropic.API
2024-03Grok-1The weights and architecture of our 314 billion parameter Mixture-of-Experts model, Grok-1.Hugging Face
2023-11MixtralOpen and portable generative AI for devs and businesses.arXivHugging Face
2023-09Baichuan 2A series of large language models developed by Baichuan Intelligent Technology.Hugging Face
2023-07GPT-4GPT-4 is OpenAI’s most advanced system, producing safer and more useful responses.API

^ Back to Contents ^

Agent

DateSourceDescriptionPaperModel
2024-10TEN AgentTEN Agent is the world’s first real-time multimodal agent integrated with the OpenAI Realtime API, RTC, and features weather checks, web search, vision, and RAG capabilities.Website
2024-08TwitterTwitter Personality is a web application that analyzes your Twitter handle to create a personalized personality profile using Wordware AI Agent.Website
2024-08MindSearch🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT).
2024-08MMRoleMMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents.arXiv
2024-08Agent KAn autoagentic AGI that is self-evolving and modular.
2024-08LangGraph StudioLangGraph Studio offers a new way to develop LLM applications by providing a specialized agent IDE that enables visualization, interaction, and debugging of complex agentic applications.
2024-07LLama Agentic SystemAgentic components of the Llama Stack APIs.
2024-07TaskGenA Task-based agentic framework building on StrictJSON outputs by LLM agents.
2024-07IoAAn open-source framework for collaborative AI agents, enabling diverse, distributed agents to team up and tackle complex tasks through internet-like connectivity.
2024-07OmAgentA multimodal agent framework for solving complex tasks.arXiv
2024-06GraphRAGA modular graph-based Retrieval-Augmented Generation (RAG) system.Website
2024-06Mixture of Agents (MoA)Mixture-of-Agents Enhances Large Language Model Capabilities.arXiv
2024-06Buffer of ThoughtsThought-Augmented Reasoning with Large Language Models.arXiv
2024-06Translation AgentAgentic translation using reflection workflow.
2024-06Atomic AgentsThe Atomic Agents framework is designed to be modular, extensible, and easy to use.
2024-05PipecatOpen Source framework for voice and multimodal conversational AI.
2024-02V-IRLGrounding Virtual Intelligence in Real Life.arXiv

^ Back to Contents ^

Audio

Audio/Text-to-Speech

DateSourceDescriptionPaperModel
2024-07CosyVoiceMulti-lingual large voice generation model, providing inference, training and deployment full-stack ability.
2024-06DEX-TTSDiffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability.arXivWebsite
2024-05ChatTTSChatTTS is a text-to-speech model designed specifically for dialogue scenario such as LLM assistant.
2023-06StyleTTS 2Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.arXivHugging Face

Audio/Automatic Speech Recognition

DateSourceDescriptionPaperModel
2024-07SenseVoiceSenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED).Hugging Face
2024-05TeleSpeech-ASRLarge speech model-super multi-dialect ASR.Hugging Face
2022-12WhisperWhisper is a general-purpose speech recognition model.arXivAPI

Audio/Audio Generation

DateSourceDescriptionPaperModel
2024-07FoleyCrafterFoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.arXivHugging Face
2024-06SEE-2-SOUNDZero-Shot Spatial Environment-to-Spatial Sound.arXiv
2024-05Make-An-Audio 3Transforming Text into Audio via Flow-based Large Diffusion Transformers.arXivHugging Face

^ Back to Contents ^

Image

DateSourceDescriptionPaperModel
2024-09StoryMakerStoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation.arXivHugging Face
2024-08CSGOCSGO: Content-Style Composition in Text-to-Image Generation.arXiv
2024-08FLUXThis repo contains minimal inference code to run text-to-image and image-to-image with our Flux latent rectified flow transformers.Hugging Face
2024-08Segment Anything Model 2 (SAM 2)SAM 2: Segment Anything in Images and Videos.arXivHugging Face
2024-07CatVTONCatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models.arXivHugging Face
2024-07UltraEditUltraEdit: Instruction-based Fine-Grained Image Editing at Scale.arXivHugging Face
2024-07UltraPixelUltraPixel: Advancing Ultra-High-Resolution Image Synthesis to New Peaks.arXiv
2024-07PaintsUndoPaintsUndo: A Base Model of Drawing Behaviors in Digital Paintings.
2024-07KolorsKolors: Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis.Hugging Face
2024-06Depth Anything V2Depth Anything V2.arXivHugging Face
2024-06AutoStudioCrafting Consistent Subjects in Multi-turn Interactive Image Generation.arXiv
2024-06MimicBrushZero-shot Image Editing with Reference Imitation.arXivHugging Face
2024-06LlamaGenAutoregressive Model Beats Diffusion: Llama for Scalable Image Generation.arXivHugging Face
2024-05OmostOmost is a project to convert LLM's coding capability to image generation (or more accurately, image composing) capability.Hugging Face
2024-05Hunyuan-DiTA Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.arXivHugging Face
2024-02MIGCMIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis.arXiv
2023-10DALL·E 3DALL·E is a AI system that can create realistic images and art from a description in natural language.API

^ Back to Contents ^

Video

DateSourceDescriptionPaperModel
2024-11LTX-VideoLTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time.Hugging Face
2024-09MIMOMIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling.arXivWebsite
2024-09DrawingSpinUpDrawingSpinUp: 3D Animation from Single Character Drawings.arXivWebsite
2024-09ViewCrafterViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis.arXivWebsite
2024-08CogVideoXCogVideoX is an open-source version of the video generation model, which is homologous to 清影.Hugging Face
2024-07ToraTora: Trajectory-oriented Diffusion Transformer for Video Generation.arXivWebsite
2024-06DiffutoonHigh-Resolution Editable Toon Shading via Diffusion Models.arXivWebsite
2024-05Video-MMEThe First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.
2024-05Video-of-ThoughtVideo-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition.Website
2024-05MOFA-VideoMOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model.arXivHugging Face
2024-05MotionLLMUnderstanding Human Behaviors from Human Motions and Videos.arXiv
2024-05ViduVidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.arXiv
2024-02SoraSora is an AI model that can create realistic and imaginative scenes from text instructions.Technical Report
2023-11PikaPika is the idea-to-video platform that sets your creativity in motion.
2023-03RunwayRunway is an applied AI research company shaping the next era of art, entertainment and human creativity.

^ Back to Contents ^

Music

DateSourceDescriptionPaperModel
2025-01YuEYuE: Open Full-song Generation Foundation Model, something similar to Suno.ai but open.Hugging Face
2024-05Diff-BGMA Diffusion Model for Video Background Music Generation.arXiv
2024-04UdioUdio - AI Music GeneratorWebsite
2023-12SunoSuno is building a future where anyone can make great music.Website
2023-12Soundry AIGenerative AI tools including text-to-sound and infinite sample packs.Website
2023-12SonautoSonauto is an AI music editor that turns prompts, lyrics, or melodies into full songs in any style.Website

^ Back to Contents ^

3D

DateSourceDescriptionPaperModel
2025-01Hunyuan3D 2.0Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation.arXivHugging Face
2024-11Hunyuan3D-1.0Hunyuan3D-1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation.arXivHugging Face
2024-093DTopia-XL3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion.arXivWebsite
2024-08ShapeSplatShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining.arXiv
2024-08SF3DSF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement.arXivHugging Face
2024-07HoloDreamerHoloDreamer: Holistic 3D Panoramic World Generation from Text Descriptions.arXivWebsite
2024-07DreamCatalystDreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation.arXivWebsite
2024-07CharacterGenCharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization.arXivWebsite
2024-07GALA3DGALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.arXivWebsite
2024-06Unique3DHigh-Quality and Efficient 3D Mesh Generation from a Single Image.arXivHugging Face
2024-06DreamGaussian4DGenerative 4D Gaussian Splatting.arXivHugging Face
2024-03GaussCtrlGaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing.arXiv
2024-03GaussianCubeA Structured and Explicit Radiance Representation for 3D Generative Modeling.arXivHugging Face
2024-03TripoSRFast 3D Object Reconstruction from a Single Image.arXivHugging Face

^ Back to Contents ^

ai
ai-agents
deeplearning-ai
llm
multi-modal
multimodal
multimodal-deep-learning

Contributors

Yuan-ManX

165 commits