Multimodal Model Compression & Optimization

33 repos

Techniques and frameworks for compressing and optimizing large multimodal models (vision-language and video-language) through distillation, efficient inference, and parameter reduction. The cluster centers on practical implementations of knowledge distillation, model pruning, and inference optimization for models that process images, video, and text together. Core repos like TimeLens, CamSFT, CamDistill, and CamInject demonstrate specific approaches to making these models smaller and faster while maintaining capability.

Python · 1
transformers ·556
safetensors ·556
video-text-to-text ·548
custom_code ·413
en ·413
multimodal ·408
video ·401
feature-extraction ·391
zh ·374
MOSS-VL ·365

ddz16/CamDistill

No description

Python

3

21 commits