ikaijua/Awesome-AIResearch

Collection of AI-related research. Welcome to submit issues and pull requests /收藏AI相关的研究,欢迎提交issues 或者pull requests

8

109 commits

updated Sep 18, 2026

See the code

README

This repo collects AI-related research and learning resources.

Spatial Intelligence

NameDescriptionLinksPublish Time
Behavior Vision SuiteBEHAVIOR Vision Suite: Customizable Dataset Generation via SimulationProject website2024

AI Agent

NameDescriptionLinksPublish Time
TEN-AgentTEN Agent is a realtime conversational AI agent powered by TEN. It seamlessly integrates the OpenAI Realtime API, RTC capabilities, and advanced features like weather updates, web search, computer vision, and Retrieval-Augmented Generation (RAG).TEN-Agent GitHub Repo stars2024
TIGER-AI-Lab/TheoremExplainAgentTheoremExplainAgent is an AI system that generates long-form Manim videos to visually explain theorems, proving its deep understanding while uncovering reasoning flaws that text alone often hides.TheoremExplainAgent GitHub Repo stars2025

Distributed Training Framework

NameDescriptionLinksPublish Time
DeepSpeedDeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.Github GitHub Repo stars-
Megatron-LMOngoing research training transformer models at scale.Github GitHub Repo stars-

Robot

NameDescriptionLinksPublish Time
RoboChallengeRoboChallenge is a real-world robotics testing and evaluation platform. Here, researchers and developers can validate and compare their robot policies in a unified environment—spanning from fundamental tasks to complex real-world scenarios.URL2025-10
huggingface/lerobotState-of-the-art Machine Learning for Real-World Robotics in PytorchGithub GitHub Repo stars2024
spirit-v1.5A Robotic Foundation Model by Spirit AIGithub GitHub Repo stars2025
TidyBotA household cleanup robot done by StanfordAILab.GitHub GitHub Repo stars2023
EurekaHuman-Level Reward Design via Coding Large Language Models, such as GPT-4, to perform in-context evolutionary optimization over reward code. Harnessing them to learn complex low-level manipulation tasks, such as dexterous pen spinningGithub GitHub Repo stars2023
NOIRNeural Signal Operated Intelligent Robots for Everyday Activities. Stanford UniversityProject website2023
robotics-survey/Awesome-Robotics-Foundation-ModelsThis repository is largely based on the following paper: Foundation Models in Robotics: Applications, Challenges, and the Future By Stanford University, Princeton University, UT Austin, NVIDIA, Scaled Foundations, Google DeepMind, TU Berlin, Shanghai Jiao Tong UniversityGithub GitHub Repo stars2023
JeffreyYH/robotics-fm-surveySurvey Paper of foundation models for robotics. paper: oward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis By CMU, Bosch Center for AI, SAIR Lab, Georgia Tech, FAIR at Meta, UC San Diego, Google DeepMindGithub GitHub Repo stars2023

Multi-modal LLM

NameDescriptionLinksPublish Time
mPLUG-DocOwlModularized Multimodal Large Language Model for Document Understanding. By Alibaba GroupGithub GitHub Repo stars2024

Vision-Language (VL) Model

NameDescriptionLinksPublish Time
DeepSeek-VLAn open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. DeepSeek-VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition, scientific literature, natural images, and embodied intelligence in complex scenarios.Github GitHub Repo stars2024
An Introduction to Vision-Language ModelingAn Introduction to Vision-Language Modeling. By Meta.URL2024
Insight-VInsight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models. By 1S-Lab, NTU. Github GitHub Repo stars2024

NeRF

NameDescriptionLinksPublish Time
NeRFCode release for NeRF (Neural Radiance Fields). Paper: https://arxiv.org/abs/2003.08934Github GitHub Repo stars2020

3D Gaussian Splatting

NameDescriptionLinksPublish Time
gaussian-splattingOriginal reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering".Github GitHub Repo stars2023

Brain Computer Interface

NameDescriptionLinksPublish Time
TBC-TJU/MetaBCIChina’s first open-source platform for non-invasive brain computer interface. The project of MetaBCI is led by Prof. Minpeng Xu from Tianjin University, China.Github GitHub Repo stars2022

LLM Datasets

NameDescriptionLinksPublish Time
Awesome-LLMs-DatasetsSummarize existing representative LLMs text datasets.Github GitHub Repo stars2024

LMMs Benchmark

NameDescriptionLinksPublish Time
SuperGPQASuperGPQA is a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines.Github GitHub Repo stars2025
mathvistaA benchmark designed to combine challenges from diverse mathematical and visual tasks. By UCLA and Microsoft ResearchProject website2023
hallucination-leaderboardLeaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents.Github GitHub Repo stars2023
GAIAA benchmark for General AI Assistants. By Meta-FAIR, Meta-GenAI, HuggingFace and AutoGPTProject website2023
microsoft/promptbenchA Unified Library for Evaluating and Understanding Large Language Models.Github GitHub Repo stars2023

Summarization

NameDescriptionLinksPublish Time
Summarization is (Almost) DeadOur findings indicate a clear preference among human evaluators for LLM-generated summaries over human-written summaries and summaries generated by fine-tuned models.https://arxiv.org/pdf/2309.09558.pdf2023

TTS

NameDescriptionLinksPublish Time
F5-TTSBy A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. By Shanghai Jiao Tong University.Github GitHub Repo stars2024
SparkAudio/Spark-TTSAn Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens. Online Demo: https://huggingface.co/spaces/Mobvoi/Offical-Spark-TTSGithub GitHub Repo stars2025
fishaudio/fish-speechBrand new TTS solution. Demo: https://fish.audio/Github GitHub Repo stars2024
VoiceCraftVoiceCraft is a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on in-the-wild data including audiobooks, internet videos, and podcasts.To clone or edit an unseen voice, VoiceCraft needs only a few seconds of reference.Github GitHub Repo stars2024
Mega-TTS 2Input text and reference audio, clone the timbre of the reference audio to generate speech corresponding to the text. By Zhejiang University and ByteDance. Paper:https://arxiv.org/abs/2307.07218URL2024
NaturalSpeech 3Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. By Microsoft Research Asia
paper: https://arxiv.org/abs/2403.03100
URL2024
BASE TTSBASE TTS is the largest TTS model to-date, trained on 100K hours of public domain speech data, achieving a new state-of-the-art in speech naturalness. By amazon.
paper:https://arxiv.org/abs/2402.08093
URL2024
metavoice-srcFoundational model for human-like, expressive TTS. Zero-shot cloning for American & British voices, with 30s reference audio.Github GitHub Repo stars2024
BarkMultilingual
Demo: https://huggingface.co/spaces/suno/bark
Paper: https://arxiv.org/abs/2209.03143
Github GitHub Repo stars2023
XTTSMultilingual
Demo: https://huggingface.co/spaces/coqui/xtts
Github GitHub Repo stars2021
OpenVoiceZH + EN
Demo: https://huggingface.co/spaces/myshell-ai/OpenVoice
Paper: https://arxiv.org/abs/2312.01479
Github GitHub Repo stars2023
TorToiSe TTSEnglish
Demo: https://huggingface.co/spaces/Manmay/tortoise-tts
Paper:https://arxiv.org/abs/2305.07243
Github GitHub Repo stars2022
GPT-SoVITSMultilingualGithub GitHub Repo stars
EmotiVoiceZH + ENGithub GitHub Repo stars2023
MeloTTShigh-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.Github GitHub Repo stars2024
Tacotron 2English
Paper: https://arxiv.org/abs/1712.05884
Unofficial Repo:Github GitHub Repo starsGDrive
SileroEM + DE + ES + EAGithub GitHub Repo stars
StyleTTS 2English
Demo: https://huggingface.co/spaces/styletts2/styletts2
Paper:https://arxiv.org/abs/2306.07691
Github GitHub Repo stars2023
AmphionDemo: https://huggingface.co/amphion
Paper: https://arxiv.org/abs/2312.09911
Github GitHub Repo stars2023
VALL-E
Paper: https://arxiv.org/abs/2301.02111
Unofficial Repo:Github GitHub Repo stars2023
PiperMultilingualGithub GitHub Repo stars
WhisperSpeechEnglish, Polish
Demo
Github GitHub Repo stars2023
HierSpeech++KR + EN
Demo:https://huggingface.co/spaces/LeeSangHoon/HierSpeech_TTS
Paper:https://arxiv.org/abs/2311.12454
Github GitHub Repo stars2023
Glow-TTSEnglish
Demo:https://jaywalnut310.github.io/glow-tts-demo/index.html
Paper:https://arxiv.org/abs/2005.11129
Github GitHub Repo stars2020
xVASynthMultilingual
Demo:https://store.steampowered.com/app/1765720/xVASynth/
Paper:https://arxiv.org/abs/2009.14153
Github GitHub Repo stars2023
IMS-ToucanMultilingual,
Demo: https://huggingface.co/spaces/Flux9665/IMS-Toucan
Paper: https://arxiv.org/abs/2206.12229
Github GitHub Repo stars2023
Matcha-TTSEnglish
Demo:https://huggingface.co/spaces/shivammehta25/Matcha-TTS
Paper:https://arxiv.org/abs/2309.03199
Repo GitHub Repo stars2023
RAD-TTSEnglish
Paper:https://openreview.net/pdf?id=0NQwnnwAORi
Github GitHub Repo stars2022
MahaTTSEnglish + Indic
Demo: Colab
Github GitHub Repo stars2023
Neural-HMM TTSEnglish
Demo:https://shivammehta25.github.io/Neural-HMM/
Paper:https://arxiv.org/abs/2108.13320
Repo GitHub Repo stars2021
pflowTTSEnglish
Paper:https://openreview.net/pdf?id=zNA7u7wtIN
Unofficial Repo GitHub Repo stars2023
PhemeEnglish
Demo:https://huggingface.co/spaces/PolyAI/pheme
Paper:https://arxiv.org/abs/2401.02839
Github GitHub Repo stars2024
TTTSZH
Demo:https://colab.research.google.com/github/adelacvg/ttts/blob/master/demo.ipynb
Github GitHub Repo stars
VITS/ MMS-TTSEnglish
Demo:https://huggingface.co/spaces/kakao-enterprise/vits
Paper:https://arxiv.org/abs/2106.06103
Github2021
OverFlow TTSEnglish
Demo:https://shivammehta25.github.io/OverFlow/
Paper: https://arxiv.org/abs/2211.06892
Github GitHub Repo stars2022

Image Generage

NameDescriptionLinksPublish Time
AnyTextMultilingual Visual Text Generation And Editing. By Alibaba GroupGithub GitHub Repo stars2023
InstantIDInstantID is a new state-of-the-art tuning-free method to achieve ID-Preserving generation with only single image, supporting various downstream tasks.Github GitHub Repo stars2023
apple/ml-mgieGuiding Instruction-based Image Editing via Multimodal Large Language Models. By Apple.Github GitHub Repo stars2024
lllyasviel/IC-LightIC-Light is a project to manipulate the illumination of images. Demo:https://huggingface.co/spaces/lllyasviel/IC-LightGithub GitHub Repo stars2024
Tencent/HunyuanDiTA Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese UnderstandingGithub GitHub Repo stars2024

Sign Language

NameDescriptionLinksPublish Time
SignLLM/Prompt2SignPrompt2Sign is first comprehensive multilingual sign language dataset, which uses tools to automate the acquisition and processing of sign language videos on the web, is an evolving data set that is efficient, lightweight, reducing the previous shortcomings.Github GitHub Repo stars2024

Video Generate

NameDescriptionLinksPublish Time
Netflix VOIDVideo Object and Interaction Deletion. A physics-aware video inpainting research that understands causal relationships and maintains scene consistency.Github GitHub Repo stars2026
Lightricks/LTX-VideoLTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time.Github
GitHub Repo stars2024
AILab-CVC/VideoGen-EvalBy Tencent. The Dawn of Video Generation: Preliminary Explorations with SORA-like ModelsGithub GitHub Repo stars2024
THUDM/CogVideoCogVideoX is an open-source version of the video generation modelGithub GitHub Repo stars2024
MusePoseMusePose is a diffusion-based and pose-guided virtual human video generation framework.By Tencent.Github GitHub Repo stars2024
ProPainterImproving Propagation and Transformer for Video Inpainting. S-Lab, Nanyang Technological UniversityGithub GitHub Repo stars2023
Emu Edit/Emu videoEmu Edit is an AI generated image model that supports modifying local content of images through text; Emu Video is an AI generated video model that also supports text modification of local content in videos.Project website2023
PixelDanceA novel approach based on diffusion models that incorporates image instructions for both the first and last frames in conjunction with text instructions for video generation. By ByteDance ResearchProject website2023
MagicDanceRealistic Human Dance Video Generation with Motions & Facial Expressions Transfer. By University of Southern CaliforniaGithub GitHub Repo stars2023
TencentARC/ MotionCtrlA Unified and Flexible Motion Controller for Video GenerationGithub GitHub Repo stars2023
DreaMovingA Human Video Generation Framework based on Diffusion Models. By Alibaba GroupGithub GitHub Repo stars2023
magicvideov2Multi-Stage High-Aesthetic Video Generation by ByteDanceURL2024
BoximatorGenerating Rich and Controllable Motions for Video Synthesis. By ByteDanceURL2024
fudan-generative-vision/champControllable and Consistent Human Image Animation with 3D Parametric GuidanceGithub GitHub Repo stars2024
TaoHuUMD/SurMoSurface-based 4D Motion Modeling for Dynamic HumanGithub GitHub Repo stars2024
ToonCrafterA research paper for generative cartoon interpolationGithub GitHub Repo stars2024

Talking Face Synthesis

NameDescriptionLinksPublish Time
INFPINFP: Audio-Driven Interactive Head Generation in Dyadic Conversations. By ByteDance.Project website2024
PersonaTalkPersonaTalk creates lip-sync visual dubbing while preserving indivisuals' talking style and facial details. Paper: https://arxiv.org/pdf/2409.05379Porject website2024
LoopyLoopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency. By Bytedance and Zhejiang UniversityProject website2024
V-ExpressV-Express aims to generate a talking head video under the control of a reference image, an audio, and a sequence of V-Kps images. By Tencent.Github GitHub Repo stars2024
InstructAvatarInstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation. By Peking UniversityProject website2024
X-LANCE/AniTalkerAnimate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingGithub GitHub Repo stars2024
VASA-1Lifelike Audio-Driven Talking Faces Generated in Real Time. By Microsoft. paper:https://arxiv.org/abs/2404.10667Project Website2024
GeneFaceGeneralized and High-Fidelity 3D Talking Face Synthesis. Zhejiang University, ByteDanceGithub GitHub Repo stars2023
GAIAZero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. GAIA (Generative AI for Avatar), which eliminates the domain priors in talking avatar generation. By MicrosoftProject Website2023

Video Comprehension

NameDescriptionLinksPublish Time
VLMEvalKitOpen-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks.Github GitHub Repo stars2024
NVlabs/VILAa multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops)Github GitHub Repo stars2024
PKU-YuanGroup/Video-LLaVAVideo-LLaVA: Learning United Visual Representation by Alignment Before ProjectionGithub GitHub Repo stars2023
evolvinglmms-lab/longvaLong Context Transfer from Language to Vision.
介绍文章:
机器之心:7B最强长视频模型! LongVA视频理解超千帧,霸榜多个榜单
Github GitHub Repo stars2024
Vision-CAIR/MiniGPT4-videoGoldfish model for long video understanding and MiniGPT4-video for short video understanding. Goldfish_websiteGithub GitHub Repo stars2024

Data Cleaning

NameDescriptionLinksPublish Time
cleanlabThe standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.Github GitHub Repo stars

3D Generate

NameDescriptionLinksPublish Time
DMV3DDenoising Multi-View Diffusion using 3D Large Reconstruction Model. A single-stage approach for high-quality text-to-3D generation and single-image reconstruction in 30s. By Adobe, Stanford, etcProject website2023
Make-A-CharacterHigh Quality Text-to-3D Character Generation within Minutes. By AlibabaGithub GitHub Repo stars2023

Object Dectection

NameDescriptionLinksPublish Time
facebookresearch/segment-anything-2Demo: https://sam2.metademolab.com/demo, blog: https://ai.meta.com/blog/segment-anything-2-video/Github GitHub Repo stars2024
open-mmlab/mmdetectionMMDetection is an open source object detection toolbox based on PyTorch.Github GitHub Repo stars
AILab-CVC/YOLO-WorldReal-Time Open-Vocabulary Object Detection. By Tencent.Github GitHub Repo stars2024
LiheYoung/Depth-AnythingDepth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation. By 1The University of Hong Kong · 2TikTok · 3Zhejiang Lab · 4Zhejiang UniversityGithub GitHub Repo stars2024
t-rexTowards Generic Object Detection via Text-Visual Prompt Synergy.Github GitHub Repo stars2024

Image/Video Enhancements

NameDescriptionLinksPublish Time
CodeFormerTowards Robust Blind Face Restoration with Codebook Lookup Transformer (NeurIPS 2022) . By S-Lab, Nanyang Technological UniversityGithub GitHub Repo stars2023

Super-Resolution

NameDescriptionLinksPublish Time
Upscale-A-VideoUpscale-A-Video is a diffusion-based model that upscales videos by taking the low-resolution video and text prompts as inputs. S-Lab, Nanyang Technological UniversityGithub GitHub Repo stars2023
ComfyUI-SUPIRSUPIR upscaling wrapper for ComfyUIGithub GitHub Repo stars2024
APISRAPISR: Anime Production Inspired Real-World Anime Super-Resolution (CVPR 2024). APISR aims at restoring and enhancing low-quality low-resolution anime images and video sources with various degradations from real-world scenarios.Github GitHub Repo stars2024
EvTextureEvent-driven Texture Enhancement for Video Super-Resolution. By University of Science and Technology of ChinaGithub GitHub Repo stars2024
jnjaby/KEEPKalman-Inspired Feature Propagation for Video Face Super-Resolution. By S-Lab, Nanyang Technological University.ECCV 2024Github GitHub Repo stars

Virtual Try-On

NameDescriptionLinksPublish Time
OutfitAnyoneOutfit Anyone: Ultra-high quality virtual try-on for Any Clothing and Any Person. Institute for Intelligent Computing, Alibaba GroupGithub GitHub Repo stars2023
OOTDiffusionOfficial implementation of OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-onGithub GitHub Repo stars Demo:https://ootd.ibot.cn/2024
ViViDViViD: Video Virtual Try-on using Diffusion Models. By AlibabaGithub GitHub Repo stars2024

AI Muisc Generation

NameDescriptionLinksPublish Time
StemGenStemGen: A music generation model that listens, ByteDance IncProject Website2023

RAG(Retrieval-Augmented Generation)

NameDescriptionLinksPublish Time
microsoft/graphragA modular graph-based Retrieval-Augmented Generation (RAG) systemGithub GitHub Repo stars2024
Retrieval-Augmented Generation for Large Language Models: A SurveyShanghai Research Institute for Intelligent Autonomous SystemsURL2023

OCR

NameDescriptionLinksPublish Time
suryaSurya is a multilingual document OCR toolkit. It can do: Accurate line-level text detectionGithub GitHub Repo stars2024
Nutlope/llama-ocrDocument to Markdown OCR library with Llama 3.2 visionGithub GitHub Repo stars2024

Visual Speech Processing

NameDescriptionLinksPublish Time
sally-sh/vsp-llmVisual Speech Processing incorporated with LLMs
paper:https://arxiv.org/abs/2402.15151v1
Github GitHub Repo stars2024

3D Human Pose Estimation

NameDescriptionLinksPublish Time
NationalGAILab/HoTHourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationGithub GitHub Repo stars2024

Computer Vision

NameDescriptionLinksPublish Time
GeneOH-DiffusionTowards Generalizable Hand-Object Interaction Denoising via Denoising DiffusionGithub GitHub Repo stars2024
Efficient-Large-Model/VILAVILA - a multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops)Github GitHub Repo stars2024
roboflow/supervisionWe write your reusable computer vision tools.Github GitHub Repo stars2023

Learning Platform

NameDescriptionLinksPublish Time
Stanford OnlineStanford's online learning platform, offering AI/ML courses and programs from Stanford University.URL-

Star History

Star 历史记录

Buy Me A Coffee

如果您喜欢这个项目,可以赞赏一下支持我们,谢谢您的支持!

Contributors

ikaijua

109 commits

ikaijua/Awesome-AIResearch

Collection of AI-related research. Welcome to submit issues and pull requests /收藏AI相关的研究,欢迎提交issues 或者pull requests

8

109 commits

updated Sep 18, 2026

See the code

README

This repo collects AI-related research and learning resources.

Spatial Intelligence

NameDescriptionLinksPublish Time
Behavior Vision SuiteBEHAVIOR Vision Suite: Customizable Dataset Generation via SimulationProject website2024

AI Agent

NameDescriptionLinksPublish Time
TEN-AgentTEN Agent is a realtime conversational AI agent powered by TEN. It seamlessly integrates the OpenAI Realtime API, RTC capabilities, and advanced features like weather updates, web search, computer vision, and Retrieval-Augmented Generation (RAG).TEN-Agent GitHub Repo stars2024
TIGER-AI-Lab/TheoremExplainAgentTheoremExplainAgent is an AI system that generates long-form Manim videos to visually explain theorems, proving its deep understanding while uncovering reasoning flaws that text alone often hides.TheoremExplainAgent GitHub Repo stars2025

Distributed Training Framework

NameDescriptionLinksPublish Time
DeepSpeedDeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.Github GitHub Repo stars-
Megatron-LMOngoing research training transformer models at scale.Github GitHub Repo stars-

Robot

NameDescriptionLinksPublish Time
RoboChallengeRoboChallenge is a real-world robotics testing and evaluation platform. Here, researchers and developers can validate and compare their robot policies in a unified environment—spanning from fundamental tasks to complex real-world scenarios.URL2025-10
huggingface/lerobotState-of-the-art Machine Learning for Real-World Robotics in PytorchGithub GitHub Repo stars2024
spirit-v1.5A Robotic Foundation Model by Spirit AIGithub GitHub Repo stars2025
TidyBotA household cleanup robot done by StanfordAILab.GitHub GitHub Repo stars2023
EurekaHuman-Level Reward Design via Coding Large Language Models, such as GPT-4, to perform in-context evolutionary optimization over reward code. Harnessing them to learn complex low-level manipulation tasks, such as dexterous pen spinningGithub GitHub Repo stars2023
NOIRNeural Signal Operated Intelligent Robots for Everyday Activities. Stanford UniversityProject website2023
robotics-survey/Awesome-Robotics-Foundation-ModelsThis repository is largely based on the following paper: Foundation Models in Robotics: Applications, Challenges, and the Future By Stanford University, Princeton University, UT Austin, NVIDIA, Scaled Foundations, Google DeepMind, TU Berlin, Shanghai Jiao Tong UniversityGithub GitHub Repo stars2023
JeffreyYH/robotics-fm-surveySurvey Paper of foundation models for robotics. paper: oward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis By CMU, Bosch Center for AI, SAIR Lab, Georgia Tech, FAIR at Meta, UC San Diego, Google DeepMindGithub GitHub Repo stars2023

Multi-modal LLM

NameDescriptionLinksPublish Time
mPLUG-DocOwlModularized Multimodal Large Language Model for Document Understanding. By Alibaba GroupGithub GitHub Repo stars2024

Vision-Language (VL) Model

NameDescriptionLinksPublish Time
DeepSeek-VLAn open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. DeepSeek-VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition, scientific literature, natural images, and embodied intelligence in complex scenarios.Github GitHub Repo stars2024
An Introduction to Vision-Language ModelingAn Introduction to Vision-Language Modeling. By Meta.URL2024
Insight-VInsight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models. By 1S-Lab, NTU. Github GitHub Repo stars2024

NeRF

NameDescriptionLinksPublish Time
NeRFCode release for NeRF (Neural Radiance Fields). Paper: https://arxiv.org/abs/2003.08934Github GitHub Repo stars2020

3D Gaussian Splatting

NameDescriptionLinksPublish Time
gaussian-splattingOriginal reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering".Github GitHub Repo stars2023

Brain Computer Interface

NameDescriptionLinksPublish Time
TBC-TJU/MetaBCIChina’s first open-source platform for non-invasive brain computer interface. The project of MetaBCI is led by Prof. Minpeng Xu from Tianjin University, China.Github GitHub Repo stars2022

LLM Datasets

NameDescriptionLinksPublish Time
Awesome-LLMs-DatasetsSummarize existing representative LLMs text datasets.Github GitHub Repo stars2024

LMMs Benchmark

NameDescriptionLinksPublish Time
SuperGPQASuperGPQA is a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines.Github GitHub Repo stars2025
mathvistaA benchmark designed to combine challenges from diverse mathematical and visual tasks. By UCLA and Microsoft ResearchProject website2023
hallucination-leaderboardLeaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents.Github GitHub Repo stars2023
GAIAA benchmark for General AI Assistants. By Meta-FAIR, Meta-GenAI, HuggingFace and AutoGPTProject website2023
microsoft/promptbenchA Unified Library for Evaluating and Understanding Large Language Models.Github GitHub Repo stars2023

Summarization

NameDescriptionLinksPublish Time
Summarization is (Almost) DeadOur findings indicate a clear preference among human evaluators for LLM-generated summaries over human-written summaries and summaries generated by fine-tuned models.https://arxiv.org/pdf/2309.09558.pdf2023

TTS

NameDescriptionLinksPublish Time
F5-TTSBy A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. By Shanghai Jiao Tong University.Github GitHub Repo stars2024
SparkAudio/Spark-TTSAn Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens. Online Demo: https://huggingface.co/spaces/Mobvoi/Offical-Spark-TTSGithub GitHub Repo stars2025
fishaudio/fish-speechBrand new TTS solution. Demo: https://fish.audio/Github GitHub Repo stars2024
VoiceCraftVoiceCraft is a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on in-the-wild data including audiobooks, internet videos, and podcasts.To clone or edit an unseen voice, VoiceCraft needs only a few seconds of reference.Github GitHub Repo stars2024
Mega-TTS 2Input text and reference audio, clone the timbre of the reference audio to generate speech corresponding to the text. By Zhejiang University and ByteDance. Paper:https://arxiv.org/abs/2307.07218URL2024
NaturalSpeech 3Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. By Microsoft Research Asia
paper: https://arxiv.org/abs/2403.03100
URL2024
BASE TTSBASE TTS is the largest TTS model to-date, trained on 100K hours of public domain speech data, achieving a new state-of-the-art in speech naturalness. By amazon.
paper:https://arxiv.org/abs/2402.08093
URL2024
metavoice-srcFoundational model for human-like, expressive TTS. Zero-shot cloning for American & British voices, with 30s reference audio.Github GitHub Repo stars2024
BarkMultilingual
Demo: https://huggingface.co/spaces/suno/bark
Paper: https://arxiv.org/abs/2209.03143
Github GitHub Repo stars2023
XTTSMultilingual
Demo: https://huggingface.co/spaces/coqui/xtts
Github GitHub Repo stars2021
OpenVoiceZH + EN
Demo: https://huggingface.co/spaces/myshell-ai/OpenVoice
Paper: https://arxiv.org/abs/2312.01479
Github GitHub Repo stars2023
TorToiSe TTSEnglish
Demo: https://huggingface.co/spaces/Manmay/tortoise-tts
Paper:https://arxiv.org/abs/2305.07243
Github GitHub Repo stars2022
GPT-SoVITSMultilingualGithub GitHub Repo stars
EmotiVoiceZH + ENGithub GitHub Repo stars2023
MeloTTShigh-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.Github GitHub Repo stars2024
Tacotron 2English
Paper: https://arxiv.org/abs/1712.05884
Unofficial Repo:Github GitHub Repo starsGDrive
SileroEM + DE + ES + EAGithub GitHub Repo stars
StyleTTS 2English
Demo: https://huggingface.co/spaces/styletts2/styletts2
Paper:https://arxiv.org/abs/2306.07691
Github GitHub Repo stars2023
AmphionDemo: https://huggingface.co/amphion
Paper: https://arxiv.org/abs/2312.09911
Github GitHub Repo stars2023
VALL-E
Paper: https://arxiv.org/abs/2301.02111
Unofficial Repo:Github GitHub Repo stars2023
PiperMultilingualGithub GitHub Repo stars
WhisperSpeechEnglish, Polish
Demo
Github GitHub Repo stars2023
HierSpeech++KR + EN
Demo:https://huggingface.co/spaces/LeeSangHoon/HierSpeech_TTS
Paper:https://arxiv.org/abs/2311.12454
Github GitHub Repo stars2023
Glow-TTSEnglish
Demo:https://jaywalnut310.github.io/glow-tts-demo/index.html
Paper:https://arxiv.org/abs/2005.11129
Github GitHub Repo stars2020
xVASynthMultilingual
Demo:https://store.steampowered.com/app/1765720/xVASynth/
Paper:https://arxiv.org/abs/2009.14153
Github GitHub Repo stars2023
IMS-ToucanMultilingual,
Demo: https://huggingface.co/spaces/Flux9665/IMS-Toucan
Paper: https://arxiv.org/abs/2206.12229
Github GitHub Repo stars2023
Matcha-TTSEnglish
Demo:https://huggingface.co/spaces/shivammehta25/Matcha-TTS
Paper:https://arxiv.org/abs/2309.03199
Repo GitHub Repo stars2023
RAD-TTSEnglish
Paper:https://openreview.net/pdf?id=0NQwnnwAORi
Github GitHub Repo stars2022
MahaTTSEnglish + Indic
Demo: Colab
Github GitHub Repo stars2023
Neural-HMM TTSEnglish
Demo:https://shivammehta25.github.io/Neural-HMM/
Paper:https://arxiv.org/abs/2108.13320
Repo GitHub Repo stars2021
pflowTTSEnglish
Paper:https://openreview.net/pdf?id=zNA7u7wtIN
Unofficial Repo GitHub Repo stars2023
PhemeEnglish
Demo:https://huggingface.co/spaces/PolyAI/pheme
Paper:https://arxiv.org/abs/2401.02839
Github GitHub Repo stars2024
TTTSZH
Demo:https://colab.research.google.com/github/adelacvg/ttts/blob/master/demo.ipynb
Github GitHub Repo stars
VITS/ MMS-TTSEnglish
Demo:https://huggingface.co/spaces/kakao-enterprise/vits
Paper:https://arxiv.org/abs/2106.06103
Github2021
OverFlow TTSEnglish
Demo:https://shivammehta25.github.io/OverFlow/
Paper: https://arxiv.org/abs/2211.06892
Github GitHub Repo stars2022

Image Generage

NameDescriptionLinksPublish Time
AnyTextMultilingual Visual Text Generation And Editing. By Alibaba GroupGithub GitHub Repo stars2023
InstantIDInstantID is a new state-of-the-art tuning-free method to achieve ID-Preserving generation with only single image, supporting various downstream tasks.Github GitHub Repo stars2023
apple/ml-mgieGuiding Instruction-based Image Editing via Multimodal Large Language Models. By Apple.Github GitHub Repo stars2024
lllyasviel/IC-LightIC-Light is a project to manipulate the illumination of images. Demo:https://huggingface.co/spaces/lllyasviel/IC-LightGithub GitHub Repo stars2024
Tencent/HunyuanDiTA Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese UnderstandingGithub GitHub Repo stars2024

Sign Language

NameDescriptionLinksPublish Time
SignLLM/Prompt2SignPrompt2Sign is first comprehensive multilingual sign language dataset, which uses tools to automate the acquisition and processing of sign language videos on the web, is an evolving data set that is efficient, lightweight, reducing the previous shortcomings.Github GitHub Repo stars2024

Video Generate

NameDescriptionLinksPublish Time
Netflix VOIDVideo Object and Interaction Deletion. A physics-aware video inpainting research that understands causal relationships and maintains scene consistency.Github GitHub Repo stars2026
Lightricks/LTX-VideoLTX-Video is the first DiT-based video generation model that can generate high-quality videos in real-time.Github
GitHub Repo stars2024
AILab-CVC/VideoGen-EvalBy Tencent. The Dawn of Video Generation: Preliminary Explorations with SORA-like ModelsGithub GitHub Repo stars2024
THUDM/CogVideoCogVideoX is an open-source version of the video generation modelGithub GitHub Repo stars2024
MusePoseMusePose is a diffusion-based and pose-guided virtual human video generation framework.By Tencent.Github GitHub Repo stars2024
ProPainterImproving Propagation and Transformer for Video Inpainting. S-Lab, Nanyang Technological UniversityGithub GitHub Repo stars2023
Emu Edit/Emu videoEmu Edit is an AI generated image model that supports modifying local content of images through text; Emu Video is an AI generated video model that also supports text modification of local content in videos.Project website2023
PixelDanceA novel approach based on diffusion models that incorporates image instructions for both the first and last frames in conjunction with text instructions for video generation. By ByteDance ResearchProject website2023
MagicDanceRealistic Human Dance Video Generation with Motions & Facial Expressions Transfer. By University of Southern CaliforniaGithub GitHub Repo stars2023
TencentARC/ MotionCtrlA Unified and Flexible Motion Controller for Video GenerationGithub GitHub Repo stars2023
DreaMovingA Human Video Generation Framework based on Diffusion Models. By Alibaba GroupGithub GitHub Repo stars2023
magicvideov2Multi-Stage High-Aesthetic Video Generation by ByteDanceURL2024
BoximatorGenerating Rich and Controllable Motions for Video Synthesis. By ByteDanceURL2024
fudan-generative-vision/champControllable and Consistent Human Image Animation with 3D Parametric GuidanceGithub GitHub Repo stars2024
TaoHuUMD/SurMoSurface-based 4D Motion Modeling for Dynamic HumanGithub GitHub Repo stars2024
ToonCrafterA research paper for generative cartoon interpolationGithub GitHub Repo stars2024

Talking Face Synthesis

NameDescriptionLinksPublish Time
INFPINFP: Audio-Driven Interactive Head Generation in Dyadic Conversations. By ByteDance.Project website2024
PersonaTalkPersonaTalk creates lip-sync visual dubbing while preserving indivisuals' talking style and facial details. Paper: https://arxiv.org/pdf/2409.05379Porject website2024
LoopyLoopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency. By Bytedance and Zhejiang UniversityProject website2024
V-ExpressV-Express aims to generate a talking head video under the control of a reference image, an audio, and a sequence of V-Kps images. By Tencent.Github GitHub Repo stars2024
InstructAvatarInstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation. By Peking UniversityProject website2024
X-LANCE/AniTalkerAnimate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingGithub GitHub Repo stars2024
VASA-1Lifelike Audio-Driven Talking Faces Generated in Real Time. By Microsoft. paper:https://arxiv.org/abs/2404.10667Project Website2024
GeneFaceGeneralized and High-Fidelity 3D Talking Face Synthesis. Zhejiang University, ByteDanceGithub GitHub Repo stars2023
GAIAZero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. GAIA (Generative AI for Avatar), which eliminates the domain priors in talking avatar generation. By MicrosoftProject Website2023

Video Comprehension

NameDescriptionLinksPublish Time
VLMEvalKitOpen-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks.Github GitHub Repo stars2024
NVlabs/VILAa multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops)Github GitHub Repo stars2024
PKU-YuanGroup/Video-LLaVAVideo-LLaVA: Learning United Visual Representation by Alignment Before ProjectionGithub GitHub Repo stars2023
evolvinglmms-lab/longvaLong Context Transfer from Language to Vision.
介绍文章:
机器之心:7B最强长视频模型! LongVA视频理解超千帧,霸榜多个榜单
Github GitHub Repo stars2024
Vision-CAIR/MiniGPT4-videoGoldfish model for long video understanding and MiniGPT4-video for short video understanding. Goldfish_websiteGithub GitHub Repo stars2024

Data Cleaning

NameDescriptionLinksPublish Time
cleanlabThe standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.Github GitHub Repo stars

3D Generate

NameDescriptionLinksPublish Time
DMV3DDenoising Multi-View Diffusion using 3D Large Reconstruction Model. A single-stage approach for high-quality text-to-3D generation and single-image reconstruction in 30s. By Adobe, Stanford, etcProject website2023
Make-A-CharacterHigh Quality Text-to-3D Character Generation within Minutes. By AlibabaGithub GitHub Repo stars2023

Object Dectection

NameDescriptionLinksPublish Time
facebookresearch/segment-anything-2Demo: https://sam2.metademolab.com/demo, blog: https://ai.meta.com/blog/segment-anything-2-video/Github GitHub Repo stars2024
open-mmlab/mmdetectionMMDetection is an open source object detection toolbox based on PyTorch.Github GitHub Repo stars
AILab-CVC/YOLO-WorldReal-Time Open-Vocabulary Object Detection. By Tencent.Github GitHub Repo stars2024
LiheYoung/Depth-AnythingDepth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation. By 1The University of Hong Kong · 2TikTok · 3Zhejiang Lab · 4Zhejiang UniversityGithub GitHub Repo stars2024
t-rexTowards Generic Object Detection via Text-Visual Prompt Synergy.Github GitHub Repo stars2024

Image/Video Enhancements

NameDescriptionLinksPublish Time
CodeFormerTowards Robust Blind Face Restoration with Codebook Lookup Transformer (NeurIPS 2022) . By S-Lab, Nanyang Technological UniversityGithub GitHub Repo stars2023

Super-Resolution

NameDescriptionLinksPublish Time
Upscale-A-VideoUpscale-A-Video is a diffusion-based model that upscales videos by taking the low-resolution video and text prompts as inputs. S-Lab, Nanyang Technological UniversityGithub GitHub Repo stars2023
ComfyUI-SUPIRSUPIR upscaling wrapper for ComfyUIGithub GitHub Repo stars2024
APISRAPISR: Anime Production Inspired Real-World Anime Super-Resolution (CVPR 2024). APISR aims at restoring and enhancing low-quality low-resolution anime images and video sources with various degradations from real-world scenarios.Github GitHub Repo stars2024
EvTextureEvent-driven Texture Enhancement for Video Super-Resolution. By University of Science and Technology of ChinaGithub GitHub Repo stars2024
jnjaby/KEEPKalman-Inspired Feature Propagation for Video Face Super-Resolution. By S-Lab, Nanyang Technological University.ECCV 2024Github GitHub Repo stars

Virtual Try-On

NameDescriptionLinksPublish Time
OutfitAnyoneOutfit Anyone: Ultra-high quality virtual try-on for Any Clothing and Any Person. Institute for Intelligent Computing, Alibaba GroupGithub GitHub Repo stars2023
OOTDiffusionOfficial implementation of OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-onGithub GitHub Repo stars Demo:https://ootd.ibot.cn/2024
ViViDViViD: Video Virtual Try-on using Diffusion Models. By AlibabaGithub GitHub Repo stars2024

AI Muisc Generation

NameDescriptionLinksPublish Time
StemGenStemGen: A music generation model that listens, ByteDance IncProject Website2023

RAG(Retrieval-Augmented Generation)

NameDescriptionLinksPublish Time
microsoft/graphragA modular graph-based Retrieval-Augmented Generation (RAG) systemGithub GitHub Repo stars2024
Retrieval-Augmented Generation for Large Language Models: A SurveyShanghai Research Institute for Intelligent Autonomous SystemsURL2023

OCR

NameDescriptionLinksPublish Time
suryaSurya is a multilingual document OCR toolkit. It can do: Accurate line-level text detectionGithub GitHub Repo stars2024
Nutlope/llama-ocrDocument to Markdown OCR library with Llama 3.2 visionGithub GitHub Repo stars2024

Visual Speech Processing

NameDescriptionLinksPublish Time
sally-sh/vsp-llmVisual Speech Processing incorporated with LLMs
paper:https://arxiv.org/abs/2402.15151v1
Github GitHub Repo stars2024

3D Human Pose Estimation

NameDescriptionLinksPublish Time
NationalGAILab/HoTHourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationGithub GitHub Repo stars2024

Computer Vision

NameDescriptionLinksPublish Time
GeneOH-DiffusionTowards Generalizable Hand-Object Interaction Denoising via Denoising DiffusionGithub GitHub Repo stars2024
Efficient-Large-Model/VILAVILA - a multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops)Github GitHub Repo stars2024
roboflow/supervisionWe write your reusable computer vision tools.Github GitHub Repo stars2023

Learning Platform

NameDescriptionLinksPublish Time
Stanford OnlineStanford's online learning platform, offering AI/ML courses and programs from Stanford University.URL-

Star History

Star 历史记录

Buy Me A Coffee

如果您喜欢这个项目,可以赞赏一下支持我们,谢谢您的支持!

Contributors

ikaijua

109 commits