27 repos
Training, fine-tuning, and evaluation frameworks for large-scale vision-language models, video understanding systems, and multimodal AI. The cluster includes foundational work on world models for embodied AI, neural rendering techniques (NeRF), video representation learning, language model alignment via DPO, and specialized fine-tuning pipelines for models like Qwen2-VL. Repositories here focus on GPU-accelerated training infrastructure, model adaptation for downstream tasks, and the intersection of vision, language, and temporal reasoning.