16 repos
Deep learning architectures and techniques for visual recognition tasks including semantic segmentation, object detection, and backbone networks. The cluster centers on Vision Transformer (ViT) variants and related foundational models, with emphasis on deformable convolutions and modern feature extraction approaches used in computer vision pipelines.