16 repos
PyTorch-based libraries and applications for vision and multimodal tasks, primarily built on HuggingFace transformers and torchvision. The cluster centers on image editing and manipulation using large vision models (notably Qwen image editors with LoRA fine-tuning), alongside optical character recognition and other vision-language model applications. Repositories here demonstrate practical implementations of state-of-the-art vision transformers and multimodal architectures for real-world computer vision workflows.