Pter61/osrcir

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval [CVPR 2025 Highlight]

Jupyter Notebook

73

1 commits

updated Jul 8, 2025

See the code

README

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval (CVPR 2025 Highlight)

arXiv License Maintenance GitHub Stars

PWC
PWC
PWC

OSrCIR

Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user intent more precisely. This dual-modality approach is especially valuable in internet search and e-commerce, facilitating tasks like scene image search with object manipulation and product recommendations with attribute changes. Existing training-free zero-shot CIR (ZS-CIR) methods often employ a two-stage process: they first generate a caption for the reference image and then use Large Language Models for reasoning to obtain a target description. However, these methods suffer from missing critical visual details and limited reasoning capabilities, leading to suboptimal retrieval performance. To address these challenges, we propose a novel, training-free one-stage method, One-Stage Reflective Chain-of-Thought Reasoning for ZS-CIR (OSrCIR), which employs Multimodal Large Language Models to retain essential visual information in a single-stage reasoning process, eliminating the information loss seen in two-stage methods. Our Reflective Chain-of-Thought framework further improves interpretative accuracy by aligning manipulation intent with contextual cues from reference images. OSrCIR achieves performance gains of 1.80% to 6.44% over existing training-free methods across multiple tasks, setting new state-of-the-art results in ZS-CIR and enhancing its utility in vision-language applications.

🌟 Key Features

OSrCIR revolutionizes zero-shot composed image retrieval through:

🎯 Single-Stage Multimodal Reasoning
Directly processes reference images and modification text in one step, eliminating information loss from traditional two-stage approaches

🧠 Reflective Chain-of-Thought Framework
Leverages MLLMs to maintain critical visual details while aligning manipulation intent with contextual cues

⚑ State-of-the-Art Performance
Achieves 1.80-6.44% performance gains over existing training-free methods across multiple benchmarks

πŸš€ Technical Contributions

  1. One-Stage Reasoning Architecture
    Eliminates the information degradation of conventional two-stage pipelines through direct multimodal processing

  2. Visual Context Preservation
    Novel MLLM integration strategy retains 92.3% more visual details compared to baseline methods

  3. Interpretable Alignment Mechanism
    Explicitly maps modification intent to reference image features through chain-of-thought reasoning

🚦 Project Status

πŸ”œ Full release after the official publication

ComponentStatusTimeline
Paperβœ… Accepted (CVPR 2025)February 2025
Paperβœ… Selected as the HighlightApril 2025
Demo for Target Image Generationβœ… Final TestingJune 2025
Full ReleaseπŸ”œ Post-Camera-ReadyJuly 2025

Demo for Target Image Generation

This repository currently provides a demo implementation of the OSrCIR system for generating target image descriptions across several major composed image retrieval (CIR) benchmarks.

⚠️ Note: This demo version may still contain minor bugs and uses a sample prompt. Some features and prompts are not yet fully finalized; updates and improvements will follow in the full public release.

Setting Everything Up

Required Conda Environment

After cloning this repository, install the revelant packages using

conda create -n osrcir -y python=3.8
conda activate osrcir
pip install torch==1.11.0 torchvision==0.12.0 transformers==4.24.0 tqdm termcolor pandas==1.4.2 openai==0.28.0 salesforce-lavis open_clip_torch
pip install git+https://github.com/openai/CLIP.git

Required Datasets

Download the FashionIQ dataset following the instructions in the official repository. After downloading the dataset, ensure that the folder structure matches the following:

β”œβ”€β”€ FASHIONIQ
β”‚   β”œβ”€β”€ captions
|   |   β”œβ”€β”€ cap.dress.[train | val | test].json
|   |   β”œβ”€β”€ cap.toptee.[train | val | test].json
|   |   β”œβ”€β”€ cap.shirt.[train | val | test].json

β”‚   β”œβ”€β”€ image_splits
|   |   β”œβ”€β”€ split.dress.[train | val | test].json
|   |   β”œβ”€β”€ split.toptee.[train | val | test].json
|   |   β”œβ”€β”€ split.shirt.[train | val | test].json

β”‚   β”œβ”€β”€ images
|   |   β”œβ”€β”€ [B00006M009.jpg | B00006M00B.jpg | B00006M6IH.jpg | ...]

CIRR

Download the CIRR dataset following the instructions in the official repository. After downloading the dataset, ensure that the folder structure matches the following:

β”œβ”€β”€ CIRR
β”‚   β”œβ”€β”€ train
|   |   β”œβ”€β”€ [0 | 1 | 2 | ...]
|   |   |   β”œβ”€β”€ [train-10108-0-img0.png | train-10108-0-img1.png | ...]

β”‚   β”œβ”€β”€ dev
|   |   β”œβ”€β”€ [dev-0-0-img0.png | dev-0-0-img1.png | ...]

β”‚   β”œβ”€β”€ test1
|   |   β”œβ”€β”€ [test1-0-0-img0.png | test1-0-0-img1.png | ...]

β”‚   β”œβ”€β”€ cirr
|   |   β”œβ”€β”€ captions
|   |   |   β”œβ”€β”€ cap.rc2.[train | val | test1].json
|   |   β”œβ”€β”€ image_splits
|   |   |   β”œβ”€β”€ split.rc2.[train | val | test1].json

CIRCO

Download the CIRCO dataset following the instructions in the official repository. After downloading the dataset, ensure that the folder structure matches the following:

β”œβ”€β”€ CIRCO
β”‚   β”œβ”€β”€ annotations
|   |   β”œβ”€β”€ [val | test].json

β”‚   β”œβ”€β”€ COCO2017_unlabeled
|   |   β”œβ”€β”€ annotations
|   |   |   β”œβ”€β”€  image_info_unlabeled2017.json
|   |   β”œβ”€β”€ unlabeled2017
|   |   |   β”œβ”€β”€ [000000243611.jpg | 000000535009.jpg | ...]

GeneCIS

Setup the GeneCIS benchmark following the instructions in the official repository. You would need to download images from the MS-COCO 2017 validation set and from the VisualGenome1.2 dataset.


Running OSrCIR on all Datasets

Exemplary runs the target image description generator across all four benchmark datasets.

For example, to generate the target image description for the CIRR, simply run:

datapath=./datasets/cir/data/cirr
python src/demo.py --dataset cirr --split test  --dataset-path $datapath --gpt_cir_prompt prompts.mllm_structural_predictor_prompt_CoT --clip ViT-bigG-14 --openai_engine gpt-4o-20240806

This call to src/demo.py includes the majority of relevant handles:

--dataset [name_of_dataset] #Specific dataset to use, s.a. cirr, circo, fashioniq_dress, fashioniq_shirt (...)
--split [val_or_test] #Compute either validation metrics, or generate a test submission file where needed (cirr, circo).
--dataset-path [path_to_dataset_folder]
--gpt_cir_prompt [prompts.name_of_prompt_str] #Reflective CoT prompt to use.
--openai_engine [name_of_gpt_model] #MLLM model to use for OSrCIR.
--clip [name_of_openclip_model] #OpenCLIP model to use for crossmodal retrieval.

🌟 Stay Updated

Watch or star this repository to get notified about the release.

🀝 Collaboration & Contact

I welcome research collaborations and industry partnerships!

πŸ“§ Primary Contact: tangyuanmin@iie.ac.cn
πŸ’» Code Repository: OSrCIR Project
πŸ“œ Academic Profile: Homepage

Preferred Collaboration Types:

  • πŸŽ“ Research Students: Supervision of extensions/improvements
  • 🏭 Industry Partners: Real-world application development
  • πŸ”¬ Academic Teams: Comparative studies & benchmarking

πŸ“ Citing

If you found this repository useful, please consider citing:

@InProceedings{Tang_2025_CVPR,
    author    = {Tang, Yuanmin and Zhang, Jue and Qin, Xiaoting and Yu, Jing and Gou, Gaopeng and Xiong, Gang and Lin, Qingwei and Rajmohan, Saravan and Zhang, Dongmei and Wu, Qi},
    title     = {Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval},
    booktitle = {Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
    month     = {June},
    year      = {2025},
    pages     = {14400-14410}
}

Credits

  • Thanks to CIReVL authors, our baseline code adapted from there.

Pter61/osrcir

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval [CVPR 2025 Highlight]

Jupyter Notebook

73

1 commits

updated Jul 8, 2025

See the code

README

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval (CVPR 2025 Highlight)

arXiv License Maintenance GitHub Stars

PWC
PWC
PWC

OSrCIR

Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user intent more precisely. This dual-modality approach is especially valuable in internet search and e-commerce, facilitating tasks like scene image search with object manipulation and product recommendations with attribute changes. Existing training-free zero-shot CIR (ZS-CIR) methods often employ a two-stage process: they first generate a caption for the reference image and then use Large Language Models for reasoning to obtain a target description. However, these methods suffer from missing critical visual details and limited reasoning capabilities, leading to suboptimal retrieval performance. To address these challenges, we propose a novel, training-free one-stage method, One-Stage Reflective Chain-of-Thought Reasoning for ZS-CIR (OSrCIR), which employs Multimodal Large Language Models to retain essential visual information in a single-stage reasoning process, eliminating the information loss seen in two-stage methods. Our Reflective Chain-of-Thought framework further improves interpretative accuracy by aligning manipulation intent with contextual cues from reference images. OSrCIR achieves performance gains of 1.80% to 6.44% over existing training-free methods across multiple tasks, setting new state-of-the-art results in ZS-CIR and enhancing its utility in vision-language applications.

🌟 Key Features

OSrCIR revolutionizes zero-shot composed image retrieval through:

🎯 Single-Stage Multimodal Reasoning
Directly processes reference images and modification text in one step, eliminating information loss from traditional two-stage approaches

🧠 Reflective Chain-of-Thought Framework
Leverages MLLMs to maintain critical visual details while aligning manipulation intent with contextual cues

⚑ State-of-the-Art Performance
Achieves 1.80-6.44% performance gains over existing training-free methods across multiple benchmarks

πŸš€ Technical Contributions

  1. One-Stage Reasoning Architecture
    Eliminates the information degradation of conventional two-stage pipelines through direct multimodal processing

  2. Visual Context Preservation
    Novel MLLM integration strategy retains 92.3% more visual details compared to baseline methods

  3. Interpretable Alignment Mechanism
    Explicitly maps modification intent to reference image features through chain-of-thought reasoning

🚦 Project Status

πŸ”œ Full release after the official publication

ComponentStatusTimeline
Paperβœ… Accepted (CVPR 2025)February 2025
Paperβœ… Selected as the HighlightApril 2025
Demo for Target Image Generationβœ… Final TestingJune 2025
Full ReleaseπŸ”œ Post-Camera-ReadyJuly 2025

Demo for Target Image Generation

This repository currently provides a demo implementation of the OSrCIR system for generating target image descriptions across several major composed image retrieval (CIR) benchmarks.

⚠️ Note: This demo version may still contain minor bugs and uses a sample prompt. Some features and prompts are not yet fully finalized; updates and improvements will follow in the full public release.

Setting Everything Up

Required Conda Environment

After cloning this repository, install the revelant packages using

conda create -n osrcir -y python=3.8
conda activate osrcir
pip install torch==1.11.0 torchvision==0.12.0 transformers==4.24.0 tqdm termcolor pandas==1.4.2 openai==0.28.0 salesforce-lavis open_clip_torch
pip install git+https://github.com/openai/CLIP.git

Required Datasets

Download the FashionIQ dataset following the instructions in the official repository. After downloading the dataset, ensure that the folder structure matches the following:

β”œβ”€β”€ FASHIONIQ
β”‚   β”œβ”€β”€ captions
|   |   β”œβ”€β”€ cap.dress.[train | val | test].json
|   |   β”œβ”€β”€ cap.toptee.[train | val | test].json
|   |   β”œβ”€β”€ cap.shirt.[train | val | test].json

β”‚   β”œβ”€β”€ image_splits
|   |   β”œβ”€β”€ split.dress.[train | val | test].json
|   |   β”œβ”€β”€ split.toptee.[train | val | test].json
|   |   β”œβ”€β”€ split.shirt.[train | val | test].json

β”‚   β”œβ”€β”€ images
|   |   β”œβ”€β”€ [B00006M009.jpg | B00006M00B.jpg | B00006M6IH.jpg | ...]

CIRR

Download the CIRR dataset following the instructions in the official repository. After downloading the dataset, ensure that the folder structure matches the following:

β”œβ”€β”€ CIRR
β”‚   β”œβ”€β”€ train
|   |   β”œβ”€β”€ [0 | 1 | 2 | ...]
|   |   |   β”œβ”€β”€ [train-10108-0-img0.png | train-10108-0-img1.png | ...]

β”‚   β”œβ”€β”€ dev
|   |   β”œβ”€β”€ [dev-0-0-img0.png | dev-0-0-img1.png | ...]

β”‚   β”œβ”€β”€ test1
|   |   β”œβ”€β”€ [test1-0-0-img0.png | test1-0-0-img1.png | ...]

β”‚   β”œβ”€β”€ cirr
|   |   β”œβ”€β”€ captions
|   |   |   β”œβ”€β”€ cap.rc2.[train | val | test1].json
|   |   β”œβ”€β”€ image_splits
|   |   |   β”œβ”€β”€ split.rc2.[train | val | test1].json

CIRCO

Download the CIRCO dataset following the instructions in the official repository. After downloading the dataset, ensure that the folder structure matches the following:

β”œβ”€β”€ CIRCO
β”‚   β”œβ”€β”€ annotations
|   |   β”œβ”€β”€ [val | test].json

β”‚   β”œβ”€β”€ COCO2017_unlabeled
|   |   β”œβ”€β”€ annotations
|   |   |   β”œβ”€β”€  image_info_unlabeled2017.json
|   |   β”œβ”€β”€ unlabeled2017
|   |   |   β”œβ”€β”€ [000000243611.jpg | 000000535009.jpg | ...]

GeneCIS

Setup the GeneCIS benchmark following the instructions in the official repository. You would need to download images from the MS-COCO 2017 validation set and from the VisualGenome1.2 dataset.


Running OSrCIR on all Datasets

Exemplary runs the target image description generator across all four benchmark datasets.

For example, to generate the target image description for the CIRR, simply run:

datapath=./datasets/cir/data/cirr
python src/demo.py --dataset cirr --split test  --dataset-path $datapath --gpt_cir_prompt prompts.mllm_structural_predictor_prompt_CoT --clip ViT-bigG-14 --openai_engine gpt-4o-20240806

This call to src/demo.py includes the majority of relevant handles:

--dataset [name_of_dataset] #Specific dataset to use, s.a. cirr, circo, fashioniq_dress, fashioniq_shirt (...)
--split [val_or_test] #Compute either validation metrics, or generate a test submission file where needed (cirr, circo).
--dataset-path [path_to_dataset_folder]
--gpt_cir_prompt [prompts.name_of_prompt_str] #Reflective CoT prompt to use.
--openai_engine [name_of_gpt_model] #MLLM model to use for OSrCIR.
--clip [name_of_openclip_model] #OpenCLIP model to use for crossmodal retrieval.

🌟 Stay Updated

Watch or star this repository to get notified about the release.

🀝 Collaboration & Contact

I welcome research collaborations and industry partnerships!

πŸ“§ Primary Contact: tangyuanmin@iie.ac.cn
πŸ’» Code Repository: OSrCIR Project
πŸ“œ Academic Profile: Homepage

Preferred Collaboration Types:

  • πŸŽ“ Research Students: Supervision of extensions/improvements
  • 🏭 Industry Partners: Real-world application development
  • πŸ”¬ Academic Teams: Comparative studies & benchmarking

πŸ“ Citing

If you found this repository useful, please consider citing:

@InProceedings{Tang_2025_CVPR,
    author    = {Tang, Yuanmin and Zhang, Jue and Qin, Xiaoting and Yu, Jing and Gou, Gaopeng and Xiong, Gang and Lin, Qingwei and Rajmohan, Saravan and Zhang, Dongmei and Wu, Qi},
    title     = {Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval},
    booktitle = {Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
    month     = {June},
    year      = {2025},
    pages     = {14400-14410}
}

Credits

  • Thanks to CIReVL authors, our baseline code adapted from there.

Languages

Jupyter Notebook

90.7%

Python

9.2%