:books: A collection of papers about Referring Image Segmentation.
830
106 commits
updated Jan 28, 2026
A collection of referring image segmentation papers and datasets.
Feel free to create a PR or an issue.

Outline
| Short name | Paper | Source | Code/Project Link |
|---|---|---|---|
| MeViS | MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions | ICCV 2023 | [dataset] [project] |
| gRefCOCO | GRES: Generalized Referring Expression Segmentation | CVPR 2023 | [dataset] [project] |
| ClevrTex | ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation | NeurIPS Datasets and Benchmarks 2021 | [project] |
| ScanRefer | ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language | ECCV 2020 | [project] |
| VGPhraseCut | PhraseCut: Language-based Image Segmentation in the Wild | CVPR 2020 | [project] |
| CLEVR-Ref+ | CLEVR-Ref+: Diagnosing Visual Reasoning with Referring Expressions | CVPR 2019 | [project] |
| UNC | Modeling context in referring expressions | ECCV 2016 | [dataset] |
| UNC+ | Modeling context in referring expressions | ECCV 2016 | [dataset] |
| Google-Ref | Generation and comprehension of unambiguous object descriptions | CVPR 2016 | [dataset] |
| ReferIt | Referit game: Referring to objects in photographs of natural scenes | EMNLP 2014 | [project] |
| Name | Workshop | Date | Submission Link |
|---|---|---|---|
| 1st MeViS Challenge | CVPR 2024 Workshop: Pixel-level Video Understanding in the Wild | May 2024 | [CodaLab] |
| RVOS Challenge | ECCV 2024 Workshop: The 6th Large-scale Video Object Segmentation Challenge | Aug 2024 | [CodaLab] |
| Short name | Paper | Source | Code/Project Link |
|---|---|---|---|
| UniPixel | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | NeurIPS 2025 | [code] [webpage] |
| PhraseClick | PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click | ECCV 2020 |
| Short name | Paper | Source | Code/Project Link |
|---|---|---|---|
| RS2-SAM2 | RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image Segmentation | AAAI 2026 | [code] |
| SurgRef | Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion | AAAI 2026 | [code] |
| RIS-LAD | RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation | AAAI 2026 | [code] |
:books: A collection of papers about Referring Image Segmentation.
830
106 commits
updated Jan 28, 2026
A collection of referring image segmentation papers and datasets.
Feel free to create a PR or an issue.

Outline
| Short name | Paper | Source | Code/Project Link |
|---|---|---|---|
| MeViS | MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions | ICCV 2023 | [dataset] [project] |
| gRefCOCO | GRES: Generalized Referring Expression Segmentation | CVPR 2023 | [dataset] [project] |
| ClevrTex | ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation | NeurIPS Datasets and Benchmarks 2021 | [project] |
| ScanRefer | ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language | ECCV 2020 | [project] |
| VGPhraseCut | PhraseCut: Language-based Image Segmentation in the Wild | CVPR 2020 | [project] |
| CLEVR-Ref+ | CLEVR-Ref+: Diagnosing Visual Reasoning with Referring Expressions | CVPR 2019 | [project] |
| UNC | Modeling context in referring expressions | ECCV 2016 | [dataset] |
| UNC+ | Modeling context in referring expressions | ECCV 2016 | [dataset] |
| Google-Ref | Generation and comprehension of unambiguous object descriptions | CVPR 2016 | [dataset] |
| ReferIt | Referit game: Referring to objects in photographs of natural scenes | EMNLP 2014 | [project] |
| Name | Workshop | Date | Submission Link |
|---|---|---|---|
| 1st MeViS Challenge | CVPR 2024 Workshop: Pixel-level Video Understanding in the Wild | May 2024 | [CodaLab] |
| RVOS Challenge | ECCV 2024 Workshop: The 6th Large-scale Video Object Segmentation Challenge | Aug 2024 | [CodaLab] |
| Short name | Paper | Source | Code/Project Link |
|---|---|---|---|
| UniPixel | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | NeurIPS 2025 | [code] [webpage] |
| PhraseClick | PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click | ECCV 2020 |
| Short name | Paper | Source | Code/Project Link |
|---|---|---|---|
| RS2-SAM2 | RS2-SAM2: Customized SAM2 for Referring Remote Sensing Image Segmentation | AAAI 2026 | [code] |
| SurgRef | Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion | AAAI 2026 | [code] |
| RIS-LAD | RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation | AAAI 2026 | [code] |