19 repos
Fine-tuning and training frameworks for vision-language models that combine image understanding with text generation capabilities. The cluster centers on the Bunny model family, which implements efficient multimodal architectures using various base language models (Phi, Qwen2, StableLM) paired with SigLIP vision encoders. Repositories here contain pretrained weights, LoRA adaptation techniques, and model configurations for building systems that understand and reason about images while generating natural language responses.