Mistral LoRA Multimodal Fine-tuning

11 repos

Fine-tuned variants of the Mistral-7B language model extended with multimodal capabilities through Low-Rank Adaptation (LoRA). These repositories combine the Mistral base model with vision encoders (CLIP, LLAVA) and audio processors (Whisper, AudioCLAP) using the PEFT framework, enabling efficient adaptation of a smaller LLM to handle image, video, and audio inputs alongside text. Projects here demonstrate practical techniques for adding multimodal reasoning to compact language models without full model retraining.

mistral-lmm ·57
finetuned ·36
multimodal ·36
peft ·32
transformers ·25
safetensors ·25
text-generation ·25
audio-text-to-text ·21
llava ·4
image-text-to-text ·4