19 repos
Deep learning frameworks and models for computer vision and multimodal tasks using PyTorch and HuggingFace Transformers. The cluster covers image processing, vision transformers, optical character recognition, and efficient model loading techniques (like LoRA fine-tuning). Most repositories are Python-based tools for building and optimizing vision models, with emphasis on practical implementations of attention mechanisms and multimodal architectures.