Image Captioning and Vision-Language Models

7 repos

Python-based tools and implementations for generating descriptive captions from images using vision-language models (VLMs). The cluster centers on JoyCaption, a prominent captioning framework, with multiple variants designed for different workflows—web UIs, CLI batch processing, and integration with ComfyUI nodes. Repositories here provide both the core captioning models and practical deployment interfaces for tasks like automated image description, dataset annotation, and creative caption generation.

Python · 5
Jupyter Notebook · 1
captioning ·1,249
joycaption ·1,249
vlm ·1,249
en ·27
safetensors ·27