10 repos
Foundation models trained to process and generate multiple modalities (text, images, video) in unified architectures, with emphasis on large-scale pretraining, instruction-tuning, and in-context learning capabilities. The cluster centers on the Emu family of models (Emu2, Emu3, Emu3.5) and their staged training pipelines, representing the state-of-practice in building generalist multimodal systems that can follow instructions and adapt to new tasks with few examples.