PyTorch Vision & Multimodal AI

19 repos

Deep learning frameworks and models for computer vision and multimodal tasks using PyTorch and HuggingFace Transformers. The cluster covers image processing, vision transformers, optical character recognition, and efficient model loading techniques (like LoRA fine-tuning). Most repositories are Python-based tools for building and optimizing vision models, with emphasis on practical implementations of attention mechanisms and multimodal architectures.

Python · 15
HTML · 2
Jupyter Notebook · 2
huggingface-transformers ·499
torchvision ·326
torch ·326
pytorch ·272
huggingface-spaces ·248
flash-attention-3 ·245
qwen-image-edit-2511 ·236
python ·200
numpy ·200
diffusers ·182