Multimodal LLM Inference & Generation

10 repos

Libraries and model implementations for vision-language models that combine text generation with image understanding capabilities. The cluster centers on the Bunny family of compact multimodal models (3B–8B parameters) optimized for efficient inference, built on foundations like Llama and leveraging transformer architectures with safetensors format for model serialization. Repositories here focus on conversational AI systems that process both text and visual inputs, relevant for practitioners building efficient multimodal applications without massive compute requirements.

transformers ·250
custom_code ·250
safetensors ·250
text-generation ·250
conversational ·182
bunny-llama ·120
bunny-phi ·91
image-text-to-text ·37
bunny-phi3 ·36
gguf ·23