Multimodal AI and Creative Generation

13 repos

Systems and frameworks for generating, processing, and synthesizing multimodal content including audio, video, text, and 3D assets using machine learning. The cluster centers on creative AI applications—from procedural audio generation and foley sound creation to video understanding and cross-modal learning. While specific language and topic metadata are sparse, the principal repositories (Uni3C, GEM, GST, TIGON, UniSH, FoleyCrafter) indicate a focus on unified multimodal architectures and generative models that bridge different data modalities for creative and synthetic content production.