14 repos
Fine-tuned and adapted implementations of LLaVA (Large Language and Vision Assistant), a multimodal model combining language and vision capabilities. These repositories focus on variants built on different base models (Llama 2, Llama 3, Vicuna) with different training approaches, compatible with Hugging Face transformers and optimized for text generation and deployment across endpoints. Explores iterative improvements like DPO (Direct Preference Optimization) and task-specific fine-tuning of vision-language understanding.