PyTorch multimodal vision models

16 repos

PyTorch-based libraries and applications for vision and multimodal tasks, primarily built on HuggingFace transformers and torchvision. The cluster centers on image editing and manipulation using large vision models (notably Qwen image editors with LoRA fine-tuning), alongside optical character recognition and other vision-language model applications. Repositories here demonstrate practical implementations of state-of-the-art vision transformers and multimodal architectures for real-world computer vision workflows.

Python · 13
Jupyter Notebook · 2
HTML · 1
huggingface-transformers ·512
torchvision ·336
torch ·336
pytorch ·282
flash-attention-3 ·253
huggingface-spaces ·248
qwen-image-edit-2511 ·244
python ·209
numpy ·209
diffusers ·191