9 repos
Video generation and diffusion-based world modeling systems, with a focus on text-to-video synthesis and temporal video understanding. The cluster centers on multimodal models that generate or predict video sequences, with particular emphasis on robotics applications and worldmodel-based approaches for understanding and generating visual dynamics. Most repos appear to be model variants or implementations exploring different scales and architectural approaches to video generation.