jianzongwu/Awesome-Open-Vocabulary

(TPAMI 2024) A Survey on Open Vocabulary Learning

1,006

119 commits

updated May 12, 2026

See the code

README

Awesome PR's Welcome

Towards Open Vocabulary Learning: A Survey

T-PAMI, 2024
Jianzong Wu * . Xiangtai Li * · Shilin Xu * · Haobo Yuan * · Henghui Ding · Yibo Yang · Xia Li · Jiangning Zhang · Yunhai Tong · Xudong Jiang · Bernard Ghanem · Dacheng Tao ·

arXiv PDF TPAMI PDF


This repo is used for recording, tracking, and benchmarking several recent open vocabulary methods to supplement our survey. If you find any work missing or have any suggestions (papers, implementations, and other resources), feel free to pull requests. We will add the missing papers to this repo as soon as possible.

🔥Add Your Paper in our Repo and Survey!!!!!

[-] You are welcome to give us an issue or PR for your open vocabulary learning work !!!!!

[-] Note that: Due to the huge paper in Arxiv, we are sorry to cover all in our survey. You can directly present a PR into this repo and we will record it for next version update of our survey.

[-] Our survey will be updated in 2024.3.

🔥New

[-] Our work is accepted by T-PAMI !!! 🔥🔥🔥

[-] We update GitHub to record the available paper by the end of 2024/1/10.

[-] We update GitHub to record the available paper by the end of 2023/7/20.

🔥Highlight!!

[1] The first survey for open vocabulary learning, including open vocabulary detection/segmentation/tracking.

[2] It also contains several related domains, including foundation model tuning and open-world detection.

[3] We list detailed results for the most representative works and give a fairer and clearer comparison of different approaches.

Introduction

This survey presents the first detailed survey on open vocabulary tasks, including open-vocabulary object detection, open-vocabulary segmentation, and 3D/video open-vocabulary tasks.

Alt Text

Summary of Contents

Methods: A Survey

Keywords

  • cap.: Use caption as auxiliary training data
  • vlm.: Use pretrained VLMs like CLIP
  • pl.: Generate pseudo labels
  • w/o ps.: Training without pixel-level supervision
  • pre.: Vision-language pretraining
  • diff.: Use diffusion models
  • unify: Unify several tasks (semantic segmentation, instance segmentation, and panoptic segmentation)
  • sam: Use SAM (Segment Anything Model)
  • open.: Demonstrated with open-set capability. (only for Video Understanding)
  • audio.: With audio modality.
  • bench: Propose a benchmark.
  • other: Other methods that cannot be grouped into above ones.
  • no-train: Does not need training.

Open Vocabulary Object Detection

YearVenueKeywordsPaper TitleCode/Project
2021CVPRcap.Open-Vocabulary Object Detection Using CaptionsCode
2022ICLRvlm.Open-vocabulary Object Detection via Vision and Language Knowledge DistillationCode
2022CVPRcap., vlm., pre.RegionCLIP: Region-based Language-Image PretrainingCode
2022CVPRvlm.Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelCode
2022CVPRvlm., cap.Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge DistillationCode
2022CVPRcap., vlm.Grounded Language-Image Pre-training[Code]
2022NeurIPScap., vlm.GLIPv2: Unifying Localization and VL UnderstandingCode
2022GCPRcap.Localized Vision-Language Matching for Open-vocabulary Object DetectionCode
2022ECCVvlm.Open-Vocabulary DETR with Conditional MatchingCode
2022ECCVvlm., cap., pl.Open Vocabulary Object Detection with Pseudo Bounding-Box LabelsCode
2022ECCVvlm.Promptdet: Towards open-vocabulary detection using uncurated imagesCode
2022ECCVvlm., pl., w/o ps.Detecting Twenty-thousand Classes using Image-level SupervisionCode
2022ECCVvlm.. pl.Exploiting unlabeled data with vision and language models for object detectionCode
2022ECCVvlm., cap.Simple Open-Vocabulary Object Detection with Vision TransformersCode
2022NeurIPSvlm., pl.Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionCode
2022NeurIPSvlm., cap.DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionN/A
2022arXivvlm.Open Vocabulary Object Detection with Proposal Mining and Prediction EqualizationCode
2022arXivvlm., pl.P3OVD: Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object DetectionN/A
2023ICLRvlm., pl.Learning Object-Language Alignments for Open-Vocabulary Object DetectionCode
2023ICLRvlm.F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language ModelsCode
2023CVPRother., vlm.Learning to Detect and Segment for Open Vocabulary Object DetectionN/A
2023CVPRvlm., cap.Aligning Bag of Regions for Open-Vocabulary Object DetectionCode
2023CVPRvlm.Object-Aware Distillation Pyramid for Open-Vocabulary Object DetectionCode
2023CVPRvlm.CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingN/A
2023CVPRvlm., pl.DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region AlignmentN/A
2023CVPRvlm.Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision TransformersN/A
2023ICMLvlm.Multi-Modal Classifiers for Open-Vocabulary Object DetectionProject
2023arXivvlm.GridCLIP: One-Stage Object Detection by Grid-Level CLIP Representation LearningN/A
2023arXivvlm., cap.Enhancing the Role of Context in Region-Word Alignment for Object DetectionN/A
2023arXivcap., pl.Open-Vocabulary Object Detection using Pseudo Caption LabelsN/A
2023arXivvlm., pl.Three ways to improve feature alignment for open vocabulary detectionN/A
2023arXivvlm.Prompt-Guided Transformers for End-to-End Open-Vocabulary Object DetectionN/A
2023TMLRvlm., cap., pl.MaMMUT: A Simple Architecture for Joint Learning for MultiModal TasksN/A
2023NeurIPSvlm., cap., pl.Scaling Open-Vocabulary Object DetectionN/A
2023arXivvlm.Open-Vocabulary Object Detection via Scene Graph DiscoveryN/A
2023ICCVvlm.Detection-Oriented Image-Text Pretraining for Open-Vocabulary DetectionCode
2023ICCVvlm.EdaDet: Open-Vocabulary Object Detection Using Early Dense AlignmentCode
2023KDDvlm.What Makes Good Open-Vocabulary Detector: A Disassembling PerspectiveN/A
2023NeurIPSvlm.CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object DetectionCode
2023arXivvlm.DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object DetectionCode
2023arXivvlm.Taming Self-Training for Open-Vocabulary Object DetectionCode
2024ICLRunify., vlm., pre.CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionCode
2023BMVCvlm.Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive OptimizationN/A
2024AAAIvlm.Simple Image-level Classification Improves Open-vocabulary Object DetectionCode
2024AAAIvlm.ProxyDet: Synthesizing Proxy Novel Classes via Classwise Mixup for Open-Vocabulary Object DetectionCode
2024AAAIunify., vlm., pre.CLIM: Contrastive Language-Image Mosaic for Region RepresentationCode
2024WACVvlm.LP-OVOD: Open-Vocabulary Object Detection by Linear ProbingCode
2024CVPRvlm.YOLO-World: Real-Time Open-Vocabulary Object DetectionCode
2024CVPRbenchThe devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understandingProject
2024ICLRvlm.LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained DescriptorsN/A
2024arXivvlm.Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary DetectionCode
2025WACVno-train, vlm., samEnhancing Novel Object Detection via Cooperative Foundational ModelsCode
2025ICRAbenchFine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and BenchmarkCode,Dataset,Benchmark

Open Vocabulary Segmentation

Semantic Segmentation

YearVenueKeywordsPaper TitleCode/Project
2022ICLRvlm.Language-driven Semantic SegmentationCode
2022CVPRcap., w/o ps.GroupViT: Semantic Segmentation Emerges from Text SupervisionCode
2022CVPRvlm.ZegFormer: Decoupling Zero-Shot Semantic SegmentationCode
2022ECCVcap., vlm.Scaling Open-Vocabulary Image Segmentation with Image-Level LabelsN/A
2022ECCVvlm, pl, w/o ps.Extract Free Dense Labels from CLIPCode
2022ECCVvlm.A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-Language ModelCode
2022ECCVvlm., cap., w/o ps.Open-world Semantic Segmentation via Contrasting and Clustering Vision-Language EmbeddingN/A
2022BMVCvlm.Open-vocabulary Semantic Segmentation with Frozen Vision-Language ModelsCode
2022arXivvlm., cap., pl, w/o ps.Perceptual Grouping in Contrastive Vision-Language ModelsCode
2022arXivvlm., cap., pl, w/o ps.SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationCode
2022arXivvlm., cap., w/o ps.Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive LearningN/A
2023CVPRvlm., pre.Generalized Decoding for Pixel, Image, and LanguageCode
2023CVPRvlm., pl.Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPCode
2023CVPRcap., vlm., w/o ps.Learning Open-vocabulary Semantic Segmentation Models From Natural Language SupervisionCode
2023CVPRvlm.Side Adapter Network for Open-Vocabulary Semantic SegmentationCodd
2023arXivvlm., unifyA Simple Framework for Open-Vocabulary Segmentation and DetectionCode
2023arXivvlm.Global Knowledge Calibration for Fast Open-Vocabulary SegmentationN/A
2023arXivvlm.CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic SegmentationCode
2023arXivvlm., unifyPrompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual RecognitionCode
2023arXivvlm., unifySegment Everything Everywhere All at OnceCode
2023arXivvlm.MVP-SEG: Multi-View Prompt Learning for Open-Vocabulary Semantic SegmentationN/A
2023arXivvlm.TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic SegmentationN/A
2023arXivvlm., w/o ps., samExploring Open-Vocabulary Semantic Segmentation without Human LabelsN/A
2023arXivvlm., unifyDaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelN/A
2023arXivdiff.Diffusion Models for Zero-Shot Open-Vocabulary SegmentationProject
2023ICCVdiff.Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion modelsProject
2023ICCVdiff.Guiding Text-to-Image Diffusion Model Towards Grounded GenerationProject
2023NeurIPScap., w/o ps.Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic SegmentationCode
2023arXivvlm.SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationCode
2023arXivvlm., no-trainPlug-and-Play, Dense-Label-Free Extraction of Open-Vocabulary Semantic Segmentation from Vision-Language ModelsN/A
2023arXivvlm., no-trainGrounding Everything: Emerging Localization Properties in Vision-Language TransformersCode
2023arXivvlm.Open-Vocabulary Segmentation with Semantic-Assisted CalibrationN/A
2023arXivvlm., no-trainSelf-Guided Open-Vocabulary Semantic SegmentationN/A
2024CVPRno-train., vlm., samCLIP as RNN: Segment Countless Visual Concepts without Training EndeavorProject
2023arXivvlm.CLIP-DINOiser: Teaching CLIP a few DINO tricksCode
2024arXivvlm., no-trainPay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic SegmentationCode
2024ECCVvlm., no-trainIn Defense of Lazy Visual Grounding for Open-Vocabulary Semantic SegmentationCode
2026AAAIcap., vlm., pre., diff.Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body PartsProject

Instance Segmentation

Panoptic Segmentation

Open Vocabulary Video Understanding

Video Classification

Tracking

YearVenueKeywordsPaper TitleCode/Project
2023CVPRvlm.,open.OVTrack: Open-Vocabulary Multiple Object TrackingProject

Video Instance Segmentation

Open Vocabulary 3D Scene Understanding

3D Classification

3D Detection

3D segmentation

Class-agnostic Detection and Segmentation

Open-World Object Detection

Open-Set Panoptic Segmentation

Acknowledgement

If you find our survey and repository useful for your research project, please consider citing our paper:

@article{wu2023open,
      title={Towards Open Vocabulary Learning: A Survey},
      author={Jianzong Wu and Xiangtai Li and Shilin Xu and Haobo Yuan and Henghui Ding and Yibo Yang and Xia Li and Jiangning Zhang and Yunhai Tong and Xudong Jiang and Bernard Ghanem and Dacheng Tao},
      year={2024},
      journal={T-PAMI},
}

Contact

jzwu@stu.pku.edu.cn
lxtpku@pku.edu.cn or xiangtai94@gmail.com

Alt Text

computer-vision
deep-learning
open-vocabulary
tpami-2024

jianzongwu/Awesome-Open-Vocabulary

(TPAMI 2024) A Survey on Open Vocabulary Learning

1,006

119 commits

updated May 12, 2026

See the code

README

Awesome PR's Welcome

Towards Open Vocabulary Learning: A Survey

T-PAMI, 2024
Jianzong Wu * . Xiangtai Li * · Shilin Xu * · Haobo Yuan * · Henghui Ding · Yibo Yang · Xia Li · Jiangning Zhang · Yunhai Tong · Xudong Jiang · Bernard Ghanem · Dacheng Tao ·

arXiv PDF TPAMI PDF


This repo is used for recording, tracking, and benchmarking several recent open vocabulary methods to supplement our survey. If you find any work missing or have any suggestions (papers, implementations, and other resources), feel free to pull requests. We will add the missing papers to this repo as soon as possible.

🔥Add Your Paper in our Repo and Survey!!!!!

[-] You are welcome to give us an issue or PR for your open vocabulary learning work !!!!!

[-] Note that: Due to the huge paper in Arxiv, we are sorry to cover all in our survey. You can directly present a PR into this repo and we will record it for next version update of our survey.

[-] Our survey will be updated in 2024.3.

🔥New

[-] Our work is accepted by T-PAMI !!! 🔥🔥🔥

[-] We update GitHub to record the available paper by the end of 2024/1/10.

[-] We update GitHub to record the available paper by the end of 2023/7/20.

🔥Highlight!!

[1] The first survey for open vocabulary learning, including open vocabulary detection/segmentation/tracking.

[2] It also contains several related domains, including foundation model tuning and open-world detection.

[3] We list detailed results for the most representative works and give a fairer and clearer comparison of different approaches.

Introduction

This survey presents the first detailed survey on open vocabulary tasks, including open-vocabulary object detection, open-vocabulary segmentation, and 3D/video open-vocabulary tasks.

Alt Text

Summary of Contents

Methods: A Survey

Keywords

  • cap.: Use caption as auxiliary training data
  • vlm.: Use pretrained VLMs like CLIP
  • pl.: Generate pseudo labels
  • w/o ps.: Training without pixel-level supervision
  • pre.: Vision-language pretraining
  • diff.: Use diffusion models
  • unify: Unify several tasks (semantic segmentation, instance segmentation, and panoptic segmentation)
  • sam: Use SAM (Segment Anything Model)
  • open.: Demonstrated with open-set capability. (only for Video Understanding)
  • audio.: With audio modality.
  • bench: Propose a benchmark.
  • other: Other methods that cannot be grouped into above ones.
  • no-train: Does not need training.

Open Vocabulary Object Detection

YearVenueKeywordsPaper TitleCode/Project
2021CVPRcap.Open-Vocabulary Object Detection Using CaptionsCode
2022ICLRvlm.Open-vocabulary Object Detection via Vision and Language Knowledge DistillationCode
2022CVPRcap., vlm., pre.RegionCLIP: Region-based Language-Image PretrainingCode
2022CVPRvlm.Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelCode
2022CVPRvlm., cap.Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge DistillationCode
2022CVPRcap., vlm.Grounded Language-Image Pre-training[Code]
2022NeurIPScap., vlm.GLIPv2: Unifying Localization and VL UnderstandingCode
2022GCPRcap.Localized Vision-Language Matching for Open-vocabulary Object DetectionCode
2022ECCVvlm.Open-Vocabulary DETR with Conditional MatchingCode
2022ECCVvlm., cap., pl.Open Vocabulary Object Detection with Pseudo Bounding-Box LabelsCode
2022ECCVvlm.Promptdet: Towards open-vocabulary detection using uncurated imagesCode
2022ECCVvlm., pl., w/o ps.Detecting Twenty-thousand Classes using Image-level SupervisionCode
2022ECCVvlm.. pl.Exploiting unlabeled data with vision and language models for object detectionCode
2022ECCVvlm., cap.Simple Open-Vocabulary Object Detection with Vision TransformersCode
2022NeurIPSvlm., pl.Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionCode
2022NeurIPSvlm., cap.DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionN/A
2022arXivvlm.Open Vocabulary Object Detection with Proposal Mining and Prediction EqualizationCode
2022arXivvlm., pl.P3OVD: Fine-grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object DetectionN/A
2023ICLRvlm., pl.Learning Object-Language Alignments for Open-Vocabulary Object DetectionCode
2023ICLRvlm.F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language ModelsCode
2023CVPRother., vlm.Learning to Detect and Segment for Open Vocabulary Object DetectionN/A
2023CVPRvlm., cap.Aligning Bag of Regions for Open-Vocabulary Object DetectionCode
2023CVPRvlm.Object-Aware Distillation Pyramid for Open-Vocabulary Object DetectionCode
2023CVPRvlm.CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingN/A
2023CVPRvlm., pl.DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region AlignmentN/A
2023CVPRvlm.Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision TransformersN/A
2023ICMLvlm.Multi-Modal Classifiers for Open-Vocabulary Object DetectionProject
2023arXivvlm.GridCLIP: One-Stage Object Detection by Grid-Level CLIP Representation LearningN/A
2023arXivvlm., cap.Enhancing the Role of Context in Region-Word Alignment for Object DetectionN/A
2023arXivcap., pl.Open-Vocabulary Object Detection using Pseudo Caption LabelsN/A
2023arXivvlm., pl.Three ways to improve feature alignment for open vocabulary detectionN/A
2023arXivvlm.Prompt-Guided Transformers for End-to-End Open-Vocabulary Object DetectionN/A
2023TMLRvlm., cap., pl.MaMMUT: A Simple Architecture for Joint Learning for MultiModal TasksN/A
2023NeurIPSvlm., cap., pl.Scaling Open-Vocabulary Object DetectionN/A
2023arXivvlm.Open-Vocabulary Object Detection via Scene Graph DiscoveryN/A
2023ICCVvlm.Detection-Oriented Image-Text Pretraining for Open-Vocabulary DetectionCode
2023ICCVvlm.EdaDet: Open-Vocabulary Object Detection Using Early Dense AlignmentCode
2023KDDvlm.What Makes Good Open-Vocabulary Detector: A Disassembling PerspectiveN/A
2023NeurIPSvlm.CoDet: Co-Occurrence Guided Region-Word Alignment for Open-Vocabulary Object DetectionCode
2023arXivvlm.DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object DetectionCode
2023arXivvlm.Taming Self-Training for Open-Vocabulary Object DetectionCode
2024ICLRunify., vlm., pre.CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionCode
2023BMVCvlm.Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive OptimizationN/A
2024AAAIvlm.Simple Image-level Classification Improves Open-vocabulary Object DetectionCode
2024AAAIvlm.ProxyDet: Synthesizing Proxy Novel Classes via Classwise Mixup for Open-Vocabulary Object DetectionCode
2024AAAIunify., vlm., pre.CLIM: Contrastive Language-Image Mosaic for Region RepresentationCode
2024WACVvlm.LP-OVOD: Open-Vocabulary Object Detection by Linear ProbingCode
2024CVPRvlm.YOLO-World: Real-Time Open-Vocabulary Object DetectionCode
2024CVPRbenchThe devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understandingProject
2024ICLRvlm.LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained DescriptorsN/A
2024arXivvlm.Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary DetectionCode
2025WACVno-train, vlm., samEnhancing Novel Object Detection via Cooperative Foundational ModelsCode
2025ICRAbenchFine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and BenchmarkCode,Dataset,Benchmark

Open Vocabulary Segmentation

Semantic Segmentation

YearVenueKeywordsPaper TitleCode/Project
2022ICLRvlm.Language-driven Semantic SegmentationCode
2022CVPRcap., w/o ps.GroupViT: Semantic Segmentation Emerges from Text SupervisionCode
2022CVPRvlm.ZegFormer: Decoupling Zero-Shot Semantic SegmentationCode
2022ECCVcap., vlm.Scaling Open-Vocabulary Image Segmentation with Image-Level LabelsN/A
2022ECCVvlm, pl, w/o ps.Extract Free Dense Labels from CLIPCode
2022ECCVvlm.A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-Language ModelCode
2022ECCVvlm., cap., w/o ps.Open-world Semantic Segmentation via Contrasting and Clustering Vision-Language EmbeddingN/A
2022BMVCvlm.Open-vocabulary Semantic Segmentation with Frozen Vision-Language ModelsCode
2022arXivvlm., cap., pl, w/o ps.Perceptual Grouping in Contrastive Vision-Language ModelsCode
2022arXivvlm., cap., pl, w/o ps.SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationCode
2022arXivvlm., cap., w/o ps.Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive LearningN/A
2023CVPRvlm., pre.Generalized Decoding for Pixel, Image, and LanguageCode
2023CVPRvlm., pl.Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPCode
2023CVPRcap., vlm., w/o ps.Learning Open-vocabulary Semantic Segmentation Models From Natural Language SupervisionCode
2023CVPRvlm.Side Adapter Network for Open-Vocabulary Semantic SegmentationCodd
2023arXivvlm., unifyA Simple Framework for Open-Vocabulary Segmentation and DetectionCode
2023arXivvlm.Global Knowledge Calibration for Fast Open-Vocabulary SegmentationN/A
2023arXivvlm.CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic SegmentationCode
2023arXivvlm., unifyPrompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual RecognitionCode
2023arXivvlm., unifySegment Everything Everywhere All at OnceCode
2023arXivvlm.MVP-SEG: Multi-View Prompt Learning for Open-Vocabulary Semantic SegmentationN/A
2023arXivvlm.TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic SegmentationN/A
2023arXivvlm., w/o ps., samExploring Open-Vocabulary Semantic Segmentation without Human LabelsN/A
2023arXivvlm., unifyDaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelN/A
2023arXivdiff.Diffusion Models for Zero-Shot Open-Vocabulary SegmentationProject
2023ICCVdiff.Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion modelsProject
2023ICCVdiff.Guiding Text-to-Image Diffusion Model Towards Grounded GenerationProject
2023NeurIPScap., w/o ps.Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic SegmentationCode
2023arXivvlm.SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationCode
2023arXivvlm., no-trainPlug-and-Play, Dense-Label-Free Extraction of Open-Vocabulary Semantic Segmentation from Vision-Language ModelsN/A
2023arXivvlm., no-trainGrounding Everything: Emerging Localization Properties in Vision-Language TransformersCode
2023arXivvlm.Open-Vocabulary Segmentation with Semantic-Assisted CalibrationN/A
2023arXivvlm., no-trainSelf-Guided Open-Vocabulary Semantic SegmentationN/A
2024CVPRno-train., vlm., samCLIP as RNN: Segment Countless Visual Concepts without Training EndeavorProject
2023arXivvlm.CLIP-DINOiser: Teaching CLIP a few DINO tricksCode
2024arXivvlm., no-trainPay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic SegmentationCode
2024ECCVvlm., no-trainIn Defense of Lazy Visual Grounding for Open-Vocabulary Semantic SegmentationCode
2026AAAIcap., vlm., pre., diff.Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body PartsProject

Instance Segmentation

Panoptic Segmentation

Open Vocabulary Video Understanding

Video Classification

Tracking

YearVenueKeywordsPaper TitleCode/Project
2023CVPRvlm.,open.OVTrack: Open-Vocabulary Multiple Object TrackingProject

Video Instance Segmentation

Open Vocabulary 3D Scene Understanding

3D Classification

3D Detection

3D segmentation

Class-agnostic Detection and Segmentation

Open-World Object Detection

Open-Set Panoptic Segmentation

Acknowledgement

If you find our survey and repository useful for your research project, please consider citing our paper:

@article{wu2023open,
      title={Towards Open Vocabulary Learning: A Survey},
      author={Jianzong Wu and Xiangtai Li and Shilin Xu and Haobo Yuan and Henghui Ding and Yibo Yang and Xia Li and Jiangning Zhang and Yunhai Tong and Xudong Jiang and Bernard Ghanem and Dacheng Tao},
      year={2024},
      journal={T-PAMI},
}

Contact

jzwu@stu.pku.edu.cn
lxtpku@pku.edu.cn or xiangtai94@gmail.com

Alt Text

computer-vision
deep-learning
open-vocabulary
tpami-2024