33 repos
Techniques and frameworks for compressing and optimizing large multimodal models (vision-language and video-language) through distillation, efficient inference, and parameter reduction. The cluster centers on practical implementations of knowledge distillation, model pruning, and inference optimization for models that process images, video, and text together. Core repos like TimeLens, CamSFT, CamDistill, and CamInject demonstrate specific approaches to making these models smaller and faster while maintaining capability.