Text-to-Video and Video Generation

109 repos across 6 sub-areas

Tools, models, and frameworks for generating videos from text descriptions and other modalities using diffusion-based and neural approaches. The cluster centers on open-source implementations of state-of-the-art video generation architectures (HunyuanVideo, CogVideoX, VideoCrafter, Flux) alongside supporting libraries for video processing, model inference, and fine-tuning. Researchers and practitioners here work with transformer-based video diffusion models, efficient video encoding schemes, and deployment pipelines for running these computationally intensive generative models.

Video Generation and Diffusion Models

37 repos

Libraries and model implementations for generating, processing, and controlling video content using diffusion-based approaches. The cluster centers on video generation frameworks built with popular diffusion libraries like Hugging Face Diffusers, leveraging safe tensor formats for model distribution. Repositories include both general-purpose video generation tools and specialized model variants (such as the Wan2.1-Fun series at different model scales) that enable video synthesis with varying levels of control and computational requirements.

Video Generation & Diffusion Models

35 repos

Diffusion-based approaches for generating and transforming video content, including text-to-video synthesis, image-to-video conversion, and video enhancement techniques. The cluster emphasizes practical implementations and fine-tuning methods (LoRAs, checkpoints) for models like VideoCrafter and CogVideoX, alongside supporting tools and reward systems for optimizing generated video quality.

Text-to-Video Generation and Benchmarking

16 repos

Tools, datasets, and evaluation frameworks for generating videos from text descriptions. This cluster focuses on benchmarking and assessing text-to-video models, with repositories providing evaluation kits, large-scale datasets, and standardized assessment methodologies. The central repositories (OpenS2V and ChronoMagic series) emphasize evaluation infrastructure and benchmark datasets for measuring video generation quality and consistency.

AI Video Generation

11 repos

Open-source models and implementations for generating videos from text descriptions and images. The cluster centers on diffusion-based approaches like CogVideoX and Flux, with a focus on making text-to-video synthesis accessible through fine-tuning, inference optimization, and custom dataset adaptation. Developers working in this area will find model weights, training pipelines, and inference frameworks primarily in Python.

Video Generation & World Models

9 repos

Video generation and diffusion-based world modeling systems, with a focus on text-to-video synthesis and temporal video understanding. The cluster centers on multimodal models that generate or predict video sequences, with particular emphasis on robotics applications and worldmodel-based approaches for understanding and generating visual dynamics. Most repos appear to be model variants or implementations exploring different scales and architectural approaches to video generation.

Video Generation from Text and Images

1 repos

Generative models and frameworks for creating videos from textual descriptions, still images, or multimodal prompts using diffusion-based and neural approaches. This cluster encompasses training pipelines, inference optimizations, and end-to-end systems for converting natural language or visual inputs into coherent video sequences, with repositories like VideoCrafter, t2v-turbo, and TIP-I2V representing different stages of the technology stack from core model development to practical deployment.