A curated, visually-organized list of Text-to-Video (T2V) products, open-source models, research papers, datasets, and benchmarks.
2026 Update: The T2V landscape has shifted dramatically. OpenAI discontinued the consumer Sora app in early 2026, while open-source models (Wan 2.2, HunyuanVideo 1.5, LTX-2.3) and commercial alternatives (Runway Gen-4.5, Seedance 2.0, Kling 3.0, Veo 3.1) now dominate.
Data collected as of June 2026. Prices, features, and availability change quickly — always check the official website before making a decision.
Key trends:
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🎬 Runway Gen-4 / Gen-4.5 | RunwayML | Professional creators | 1080p (4K upscale) | @reference consistency, world-class physics, Aleph editing | runwayml.com |
| 🐉 Kling 3.0 | Kuaishou | Motion-heavy / cinematic | 1080p, 15 s | Best-in-class motion, generous free tier | klingai.com |
| 🔍 Veo 3 / Veo 3.1 | Google DeepMind | 4K broadcast production | Native 4K | Scene extension, native audio + lip-sync | deepmind.google |
| 🌱 Seedance 2.0 | ByteDance | Multi-shot storytelling | 2K, 60 s multi-shot | 12 mixed inputs, native audio-video joint gen | seedance.tv |
| ✨ Luma Dream Machine | Luma AI | Action / sports / physics | 4K | Realistic motion blur, fluid dynamics | lumalabs.ai |
| 🎭 Pika 2.5 / 3.0 | Pika Labs | Social / stylized content | 2K | Fast, cheap, strong style transfer | pika.art |
| 🌀 Hailuo AI | MiniMax | Realistic humans / prompt adherence | 1080p, 10 s | Strong physical realism, #1 in China | hailuoai.video |
| 🌀 MiniMax H3 (Third-Party) | MiniMax3.org (independent) | Cinematic text/image/video/audio-reference generation | 2K | 768p/2K output and native audio; third-party platform, not an official MiniMax product | minimax3.org |
| 🎬 PixVerse V6 | PixVerse | Anime / stylized content | 1080p, 15 s | Character consistency engine, 20+ camera controls, native audio | pixverse.ai |
| 🎥 Vidu Q1/Q2 | Shengshu / Tsinghua | Highly consistent T2V | 1080p, 16 s | U-ViT backbone, subject consistency, 1080p generation | vidu.com |
| 🔬 Lumiere | Google DeepMind | Research T2V / I2V / editing | 720p | Space-time U-Net, single-pass temporal generation | lumiere-video.github.io |
| 🎞️ Seele TV | Seele AI | Cinematic sequence workflows | Varies by selected model | Visual references, shot-level camera direction, and continuity-oriented video creation | seele.tv |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🧑💼 Synthesia | Synthesia | Corporate training / avatars | 1080p | 100+ avatars, 130+ languages | synthesia.io |
| 🎤 DeepBrain AI | DeepBrain AI | Hyper-realistic avatars | 1080p | PPT-to-video, chroma key, native AI anchors | deepbrain.io |
| 🎥 HeyGen | HeyGen | AI avatars / translation | 1080p | 100+ avatars, multilingual voice clone, video translation | heygen.com |
| 🏢 Colossyan | Colossyan | Corporate training / e-learning | 1080p | SCORM export, branching quizzes, dual-avatar scenes | colossyan.com |
| ⚡ Elai.io | Elai | Document-to-video automation | 1080p | URL/PPTX-to-video, 100+ languages, ~1.5 min rendering | elai.io |
| 🖼️ D-ID | D-ID | Talking photos / short outreach | 1080p | Animate any photo, live streaming API, cheapest entry | d-id.com |
| 📺 Hour One | Hour One | Enterprise presentations | 1080p | Broadcast-ready virtual presenters, clean UI | hourone.ai |
| 🧑🏫 Vidnoz AI | Vidnoz | Budget avatar videos / training | 1080p | Cheap avatar generation, multilingual voiceovers | vidnoz.com |
| 🎨 Steve.AI | Steve.AI | Animated explainer videos | 4K | Multiple animation styles, character library | steve.ai |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🛠️ Krea AI | Krea | Model-agnostic access | Varies | Unified UI for Kling, Hailuo, Luma, Runway, Pika | krea.ai |
| 🌀 Morphic / Morph Studio | Morph Studio | All-in-one creative platform | Varies | 15+ models, canvas editing, custom model training | morphic.com |
| ✂️ InVideo AI | InVideo | Social media / marketing | 1080p | 5000+ templates, AI script-to-video workflow | invideo.io |
| 🎬 YumCut | YumCut | Self-hosted vertical shorts | 9:16 | Script, voice, visuals, captions, and automation API | yumcut.com · GitHub |
| 🎩 Magic Hour | Magic Hour | Multi-format creative suite | 1080p | Face swap, talking photos, headshots, clothes swapper | magichour.ai |
| 🛒 Creatify | Creatify | UGC-style ad generation | 1080p | E-commerce focused, ad performance tracking | creatify.ai |
| 🎞️ Vivideo | Vivideo | Model-agnostic short-video creation | Varies | Unified access to multiple T2V/I2V models, synced audio, text- and image-to-video | vivideo.ai |
| 🎞️ cv.cm/v (Cloud Clipboard AI Studio) | Cloud Clipboard | Seedance-based video workflows | Varies | Queue-free Seedance 2.0, image generation, canvas, and short-drama agent | cv.cm/v |
| 🎞️ Pixo | Pixo | Story-driven / storyboard-to-film | Varies | Story idea → storyboard → scene-by-scene generation → finished video, agent-assisted editing | pixo.video |
| 📣 Advibly | Advibly | On-brand ad creative (UGC / image / video) | Varies | Brand + product context drives generation; UGC and talking-actor ads, carousels, voiceover and music; agent-native via MCP | advibly.com |
| 📝 videos.social | videos.social | Content marketers / faceless drafts | Varies | Editable draft from blogs, PDFs, and prompts (script, scenes, voiceover); 1 free render; 1 credit = 1 render | videos.social |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| ⚠️ | Consumer app shut down March 2026 | ||||
| ⚠️ | Consumer product shut down in 2026 |
Runway Gen-4.5Ranked #1 on Artificial Analysis T2V benchmark (1,247 Elo). Native text-to-video with character-locking |
Kling 3.0Best-in-class motion quality, 66 free credits/day, and strong temporal consistency for cinematic content. Try Kling → |
Veo 3.1First true native 4K T2V model with scene extension to 60+ seconds and synchronized lip-sync audio. Try Veo → |
Seedance 2.0ByteDance's multi-shot native audio-video generator. Up to 12 mixed inputs, 60 fps, timeline prompting. Try Seedance → |
ColossyanSCORM-compliant corporate training with branching quizzes, dual-avatar conversations, and 80+ languages. Try Colossyan → |
Magic HourCreative multi-format suite: face swap, talking photos, headshots, and clothes swapper for social content. Try Magic Hour → |
PixVerse V6Specialized in anime and stylized content with character consistency, 20+ camera controls, and native audio. Try PixVerse → |
Morphic / Morph StudioAll-in-one creative platform with 15+ models, canvas-based editing, and custom model training for teams. Try Morphic → |
| Tool | What it does | Link |
|---|---|---|
| Podframes | Generates two-host AI podcast videos with scripted dialogue, voices, lip-sync, and captions. | github.com/Jellypod-Inc/podframes |
| TubePrompter | Converts existing videos into optimized text-to-video prompts for Sora, Veo, Runway, etc. | tubeprompter.com |
| Vadoo AI | AI shorts automation platform for faceless channels and social clips. | vadoo.tv |
| Omni-Rewriter | Open agentic prompt-expansion harness for image/video model dialects (schema + validation + bounded repair; expand ≠ generate). | github.com/WayneJin0918/Omni-Rewriter |
| BeatDesign | Open-source, local-first AI media workbench combining a Canvas, short-form video editor, shared Assets, and MCP control for image and video workflows. | github.com/BeatAPI/BeatDesign |
| minimax-h3-1000-prompts | Curated index of the MiniMax H3 1K prompt dataset: 3-field prompt anatomy, 10 hand-picked reusable prompts, and a model comparison. Interactive atlas of all 1,000 clips. | github.com/yangzhou-chaofan/minimax-h3-1000-prompts · neta.art atlas |
| Your Goal | Recommended Tool | Why |
|---|---|---|
| Professional filmmaking / VFX | Runway Gen-4.5 | Best creative control, reference consistency, Aleph editing |
| Best motion realism on a budget | Kling 3.0 | Top-tier motion, generous free tier, native audio |
| 4K broadcast / cinematic scenes | Veo 3.1 | Native 4K, scene extension, lip-sync audio |
| Multi-shot storytelling | Seedance 2.0 | Unified audio-video generation, timeline prompting |
| Anime / stylized content | PixVerse V6 | Character consistency, camera controls, native audio |
| Corporate training / LMS | Synthesia / Colossyan | SCORM, avatars, quizzes, multilingual |
| Avatar marketing videos | HeyGen / D-ID | Expressive avatars, fast generation, API access |
| Real-time creative exploration | Krea AI | 64+ models, sub-50ms feedback |
| Local / open-source deployment | Wan 2.2 / HunyuanVideo 1.5 | Apache 2.0; Wan 2.2 offers a 5B model for 24 GB GPUs |
Get started with the most popular open-source models:
Wan 2.2
git clone https://github.com/Wan-Video/Wan2.2
cd Wan2.2
pip install -r requirements.txt
HunyuanVideo 1.5
git clone https://github.com/Tencent-Hunyuan/HunyuanVideo
cd HunyuanVideo
# Follow the official guide for sampling and I2V inference
CogVideoX
git clone https://github.com/THUDM/CogVideo
cd CogVideo
pip install -r requirements.txt
For detailed inference configs, LoRA fine-tuning, and ComfyUI workflows, see each repository's README.
| Dataset | Scale | Highlights | Link |
|---|---|---|---|
| WebVid-10M | 10.7M clips | Large-scale text-video pairs scraped from the web | m-bain/webvid |
| InternVid | 7M+ clips | High-quality video-text dataset with multimodal annotations | OpenGVLab/InternVid |
| HD-VILA-100M | 100M clips | High-resolution long-form videos with dense captions | microsoft/HQD |
| Panda-70M | 70M clips | Large dataset of high-quality video-caption pairs | snap-research/Panda-70M |
| Vript | 420K clips | Ultra-detailed captions (~145 words), shot type and camera movement | Vript |
| MiraData | 798K clips | Long-duration videos (avg 72s), 1080p, 318-word captions | MiraData |
| OpenVid-1M | 1M clips | High-quality, diverse scenarios, ~98-word captions | OpenVid-1M |
| CI-VID | 1M clips / 717K seqs | Coherent interleaved text-video dataset | CI-VID |
| HD-VG-130M | 130M clips | High-definition, watermark-free video-text pairs | HD-VG-130M |
| VidProM | — | Large prompt-gallery dataset for video generation | VidProM |
| Benchmark | Focus | Link |
|---|---|---|
| VBench / VBench-2.0 | Comprehensive video generation evaluation suite | Vchitect/VBench |
| T2V-CompBench | Compositional text-to-video generation | T2V-CompBench |
| T2VTextBench | Human evaluation of textual control in video generation | T2VTextBench |
| ChronoMagic-Bench | Metamorphic / time-lapse text-to-video evaluation | ChronoMagic-Bench |
| BRITE | Reliable T2V evaluation on implausible scenarios | arXiv:2605.00873 |
| VideoEval | Low-cost evaluation of video foundation models | VideoEval |
| VideoScore2 | Think-before-you-score reward model / metric for generated video: visual quality, T2V alignment & physical consistency with chain-of-thought (ships VideoScore-Bench-v2) | Paper · Code |
Contributions are welcome! Please open a pull request to add new products, papers, models, datasets, or benchmarks. Keep entries concise and include both a paper link and a code/project link when available.
If you have any questions, feel free to contact jianzhnie.
Awesome-Text-To-Video is released under the Apache 2.0 license.
This list was compiled and updated using information from the following sources.
Python
52.7%
Shell
47.3%
A curated, visually-organized list of Text-to-Video (T2V) products, open-source models, research papers, datasets, and benchmarks.
2026 Update: The T2V landscape has shifted dramatically. OpenAI discontinued the consumer Sora app in early 2026, while open-source models (Wan 2.2, HunyuanVideo 1.5, LTX-2.3) and commercial alternatives (Runway Gen-4.5, Seedance 2.0, Kling 3.0, Veo 3.1) now dominate.
Data collected as of June 2026. Prices, features, and availability change quickly — always check the official website before making a decision.
Key trends:
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🎬 Runway Gen-4 / Gen-4.5 | RunwayML | Professional creators | 1080p (4K upscale) | @reference consistency, world-class physics, Aleph editing | runwayml.com |
| 🐉 Kling 3.0 | Kuaishou | Motion-heavy / cinematic | 1080p, 15 s | Best-in-class motion, generous free tier | klingai.com |
| 🔍 Veo 3 / Veo 3.1 | Google DeepMind | 4K broadcast production | Native 4K | Scene extension, native audio + lip-sync | deepmind.google |
| 🌱 Seedance 2.0 | ByteDance | Multi-shot storytelling | 2K, 60 s multi-shot | 12 mixed inputs, native audio-video joint gen | seedance.tv |
| ✨ Luma Dream Machine | Luma AI | Action / sports / physics | 4K | Realistic motion blur, fluid dynamics | lumalabs.ai |
| 🎭 Pika 2.5 / 3.0 | Pika Labs | Social / stylized content | 2K | Fast, cheap, strong style transfer | pika.art |
| 🌀 Hailuo AI | MiniMax | Realistic humans / prompt adherence | 1080p, 10 s | Strong physical realism, #1 in China | hailuoai.video |
| 🌀 MiniMax H3 (Third-Party) | MiniMax3.org (independent) | Cinematic text/image/video/audio-reference generation | 2K | 768p/2K output and native audio; third-party platform, not an official MiniMax product | minimax3.org |
| 🎬 PixVerse V6 | PixVerse | Anime / stylized content | 1080p, 15 s | Character consistency engine, 20+ camera controls, native audio | pixverse.ai |
| 🎥 Vidu Q1/Q2 | Shengshu / Tsinghua | Highly consistent T2V | 1080p, 16 s | U-ViT backbone, subject consistency, 1080p generation | vidu.com |
| 🔬 Lumiere | Google DeepMind | Research T2V / I2V / editing | 720p | Space-time U-Net, single-pass temporal generation | lumiere-video.github.io |
| 🎞️ Seele TV | Seele AI | Cinematic sequence workflows | Varies by selected model | Visual references, shot-level camera direction, and continuity-oriented video creation | seele.tv |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🧑💼 Synthesia | Synthesia | Corporate training / avatars | 1080p | 100+ avatars, 130+ languages | synthesia.io |
| 🎤 DeepBrain AI | DeepBrain AI | Hyper-realistic avatars | 1080p | PPT-to-video, chroma key, native AI anchors | deepbrain.io |
| 🎥 HeyGen | HeyGen | AI avatars / translation | 1080p | 100+ avatars, multilingual voice clone, video translation | heygen.com |
| 🏢 Colossyan | Colossyan | Corporate training / e-learning | 1080p | SCORM export, branching quizzes, dual-avatar scenes | colossyan.com |
| ⚡ Elai.io | Elai | Document-to-video automation | 1080p | URL/PPTX-to-video, 100+ languages, ~1.5 min rendering | elai.io |
| 🖼️ D-ID | D-ID | Talking photos / short outreach | 1080p | Animate any photo, live streaming API, cheapest entry | d-id.com |
| 📺 Hour One | Hour One | Enterprise presentations | 1080p | Broadcast-ready virtual presenters, clean UI | hourone.ai |
| 🧑🏫 Vidnoz AI | Vidnoz | Budget avatar videos / training | 1080p | Cheap avatar generation, multilingual voiceovers | vidnoz.com |
| 🎨 Steve.AI | Steve.AI | Animated explainer videos | 4K | Multiple animation styles, character library | steve.ai |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| 🛠️ Krea AI | Krea | Model-agnostic access | Varies | Unified UI for Kling, Hailuo, Luma, Runway, Pika | krea.ai |
| 🌀 Morphic / Morph Studio | Morph Studio | All-in-one creative platform | Varies | 15+ models, canvas editing, custom model training | morphic.com |
| ✂️ InVideo AI | InVideo | Social media / marketing | 1080p | 5000+ templates, AI script-to-video workflow | invideo.io |
| 🎬 YumCut | YumCut | Self-hosted vertical shorts | 9:16 | Script, voice, visuals, captions, and automation API | yumcut.com · GitHub |
| 🎩 Magic Hour | Magic Hour | Multi-format creative suite | 1080p | Face swap, talking photos, headshots, clothes swapper | magichour.ai |
| 🛒 Creatify | Creatify | UGC-style ad generation | 1080p | E-commerce focused, ad performance tracking | creatify.ai |
| 🎞️ Vivideo | Vivideo | Model-agnostic short-video creation | Varies | Unified access to multiple T2V/I2V models, synced audio, text- and image-to-video | vivideo.ai |
| 🎞️ cv.cm/v (Cloud Clipboard AI Studio) | Cloud Clipboard | Seedance-based video workflows | Varies | Queue-free Seedance 2.0, image generation, canvas, and short-drama agent | cv.cm/v |
| 🎞️ Pixo | Pixo | Story-driven / storyboard-to-film | Varies | Story idea → storyboard → scene-by-scene generation → finished video, agent-assisted editing | pixo.video |
| 📣 Advibly | Advibly | On-brand ad creative (UGC / image / video) | Varies | Brand + product context drives generation; UGC and talking-actor ads, carousels, voiceover and music; agent-native via MCP | advibly.com |
| 📝 videos.social | videos.social | Content marketers / faceless drafts | Varies | Editable draft from blogs, PDFs, and prompts (script, scenes, voiceover); 1 free render; 1 credit = 1 render | videos.social |
| Product | Maker | Best For | Max Output | Highlights | Link |
|---|---|---|---|---|---|
| ⚠️ | Consumer app shut down March 2026 | ||||
| ⚠️ | Consumer product shut down in 2026 |
Runway Gen-4.5Ranked #1 on Artificial Analysis T2V benchmark (1,247 Elo). Native text-to-video with character-locking |
Kling 3.0Best-in-class motion quality, 66 free credits/day, and strong temporal consistency for cinematic content. Try Kling → |
Veo 3.1First true native 4K T2V model with scene extension to 60+ seconds and synchronized lip-sync audio. Try Veo → |
Seedance 2.0ByteDance's multi-shot native audio-video generator. Up to 12 mixed inputs, 60 fps, timeline prompting. Try Seedance → |
ColossyanSCORM-compliant corporate training with branching quizzes, dual-avatar conversations, and 80+ languages. Try Colossyan → |
Magic HourCreative multi-format suite: face swap, talking photos, headshots, and clothes swapper for social content. Try Magic Hour → |
PixVerse V6Specialized in anime and stylized content with character consistency, 20+ camera controls, and native audio. Try PixVerse → |
Morphic / Morph StudioAll-in-one creative platform with 15+ models, canvas-based editing, and custom model training for teams. Try Morphic → |
| Tool | What it does | Link |
|---|---|---|
| Podframes | Generates two-host AI podcast videos with scripted dialogue, voices, lip-sync, and captions. | github.com/Jellypod-Inc/podframes |
| TubePrompter | Converts existing videos into optimized text-to-video prompts for Sora, Veo, Runway, etc. | tubeprompter.com |
| Vadoo AI | AI shorts automation platform for faceless channels and social clips. | vadoo.tv |
| Omni-Rewriter | Open agentic prompt-expansion harness for image/video model dialects (schema + validation + bounded repair; expand ≠ generate). | github.com/WayneJin0918/Omni-Rewriter |
| BeatDesign | Open-source, local-first AI media workbench combining a Canvas, short-form video editor, shared Assets, and MCP control for image and video workflows. | github.com/BeatAPI/BeatDesign |
| minimax-h3-1000-prompts | Curated index of the MiniMax H3 1K prompt dataset: 3-field prompt anatomy, 10 hand-picked reusable prompts, and a model comparison. Interactive atlas of all 1,000 clips. | github.com/yangzhou-chaofan/minimax-h3-1000-prompts · neta.art atlas |
| Your Goal | Recommended Tool | Why |
|---|---|---|
| Professional filmmaking / VFX | Runway Gen-4.5 | Best creative control, reference consistency, Aleph editing |
| Best motion realism on a budget | Kling 3.0 | Top-tier motion, generous free tier, native audio |
| 4K broadcast / cinematic scenes | Veo 3.1 | Native 4K, scene extension, lip-sync audio |
| Multi-shot storytelling | Seedance 2.0 | Unified audio-video generation, timeline prompting |
| Anime / stylized content | PixVerse V6 | Character consistency, camera controls, native audio |
| Corporate training / LMS | Synthesia / Colossyan | SCORM, avatars, quizzes, multilingual |
| Avatar marketing videos | HeyGen / D-ID | Expressive avatars, fast generation, API access |
| Real-time creative exploration | Krea AI | 64+ models, sub-50ms feedback |
| Local / open-source deployment | Wan 2.2 / HunyuanVideo 1.5 | Apache 2.0; Wan 2.2 offers a 5B model for 24 GB GPUs |
Get started with the most popular open-source models:
Wan 2.2
git clone https://github.com/Wan-Video/Wan2.2
cd Wan2.2
pip install -r requirements.txt
HunyuanVideo 1.5
git clone https://github.com/Tencent-Hunyuan/HunyuanVideo
cd HunyuanVideo
# Follow the official guide for sampling and I2V inference
CogVideoX
git clone https://github.com/THUDM/CogVideo
cd CogVideo
pip install -r requirements.txt
For detailed inference configs, LoRA fine-tuning, and ComfyUI workflows, see each repository's README.
| Dataset | Scale | Highlights | Link |
|---|---|---|---|
| WebVid-10M | 10.7M clips | Large-scale text-video pairs scraped from the web | m-bain/webvid |
| InternVid | 7M+ clips | High-quality video-text dataset with multimodal annotations | OpenGVLab/InternVid |
| HD-VILA-100M | 100M clips | High-resolution long-form videos with dense captions | microsoft/HQD |
| Panda-70M | 70M clips | Large dataset of high-quality video-caption pairs | snap-research/Panda-70M |
| Vript | 420K clips | Ultra-detailed captions (~145 words), shot type and camera movement | Vript |
| MiraData | 798K clips | Long-duration videos (avg 72s), 1080p, 318-word captions | MiraData |
| OpenVid-1M | 1M clips | High-quality, diverse scenarios, ~98-word captions | OpenVid-1M |
| CI-VID | 1M clips / 717K seqs | Coherent interleaved text-video dataset | CI-VID |
| HD-VG-130M | 130M clips | High-definition, watermark-free video-text pairs | HD-VG-130M |
| VidProM | — | Large prompt-gallery dataset for video generation | VidProM |
| Benchmark | Focus | Link |
|---|---|---|
| VBench / VBench-2.0 | Comprehensive video generation evaluation suite | Vchitect/VBench |
| T2V-CompBench | Compositional text-to-video generation | T2V-CompBench |
| T2VTextBench | Human evaluation of textual control in video generation | T2VTextBench |
| ChronoMagic-Bench | Metamorphic / time-lapse text-to-video evaluation | ChronoMagic-Bench |
| BRITE | Reliable T2V evaluation on implausible scenarios | arXiv:2605.00873 |
| VideoEval | Low-cost evaluation of video foundation models | VideoEval |
| VideoScore2 | Think-before-you-score reward model / metric for generated video: visual quality, T2V alignment & physical consistency with chain-of-thought (ships VideoScore-Bench-v2) | Paper · Code |
Contributions are welcome! Please open a pull request to add new products, papers, models, datasets, or benchmarks. Keep entries concise and include both a paper link and a code/project link when available.
If you have any questions, feel free to contact jianzhnie.
Awesome-Text-To-Video is released under the Apache 2.0 license.
This list was compiled and updated using information from the following sources.
Python
52.7%
Shell
47.3%