[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
See the codeJun. 17, 2025 🔥 We have released the checkpoints of our fine-tuned model.Apr. 13, 2024 We released the SPEC dataset and the code for evaluation, sorry for the delay :relaxed:.Feb. 28, 2024 Our work has been accepted by CVPR 2024 :tada:.To evaluate the understanding capability of visual-language models on fine-grained concepts, we propose a new benchmark, SPEC, which consists of six distinct subsets, distributed across the dimensions of Size, Position, Existence, and Count. Each test case consists of an image candidate set, which differs only in certain visual concepts, and a text candidate set, which differs only in the corresponding language concept.
git clone https://github.com/wjpoom/SPEC.git
cd SPEC/
pip install -e .
/path/to/save/data with a specified dir to store the data.import zipfile
import os
from huggingface_hub import hf_hub_download
data_root = '/path/to/save/data'
hf_hub_download(repo_id='wjpoom/SPEC', repo_type='dataset', filename='data.zip', local_dir=data_root)
with zipfile.ZipFile(os.path.join(data_root, 'data.zip'), 'r') as zip_ref:
zip_ref.extractall(os.path.join(data_root))
os.remove(os.path.join(data_root, 'data.zip'))
pip install open_clip_torch
mkdir checkpoints
huggingface-cli download wjpoom/SPEC-CLIP-ViT-B-32 --local-dir checkpoints/SPEC-CLIP-ViT-B-32
import torch
from PIL import Image
import open_clip
model, _, preprocess = open_clip.create_model_and_transforms('ViT-B-32', pretrained='checkpoints/SPEC-CLIP-ViT-B-32', load_weights_only=False)
model.eval()
tokenizer = open_clip.get_tokenizer('ViT-B-32')
image = preprocess(Image.open("assets/image.png")).unsqueeze(0)
text = tokenizer([
"the broccoli is situated above the backpack.",
"the broccoli is situated to the right of the backpack",
"the broccoli is positioned on the left of the backpack.",
"the broccoli is placed beneath the backpack."
])
with torch.no_grad(), torch.autocast("cuda"):
image_features = model.encode_image(image)
text_features = model.encode_text(text)
image_features /= image_features.norm(dim=-1, keepdim=True)
text_features /= text_features.norm(dim=-1, keepdim=True)
text_probs = (100.0 * image_features @ text_features.T).softmax(dim=-1)
print("Label probs:", text_probs)
Part of this repository is built upon ARO, thanks for the well-organized codebase.
Feel free to contact us if you have any questions or suggestions
Email (Wujian Peng): wjpeng24@m.fudan.edu.cn
If you use our code or data in this repo or find our work helpful, please consider giving a citation:
@inproceedings{peng2024synthesize,
title={Synthesize diagnose and optimize: Towards fine-grained vision-language understanding},
author={Peng, Wujian and Xie, Sicheng and You, Zuyao and Lan, Shiyi and Wu, Zuxuan},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={13279--13288},
year={2024}
}
Jupyter Notebook
98.6%
Python
1.4%
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
See the codeJun. 17, 2025 🔥 We have released the checkpoints of our fine-tuned model.Apr. 13, 2024 We released the SPEC dataset and the code for evaluation, sorry for the delay :relaxed:.Feb. 28, 2024 Our work has been accepted by CVPR 2024 :tada:.To evaluate the understanding capability of visual-language models on fine-grained concepts, we propose a new benchmark, SPEC, which consists of six distinct subsets, distributed across the dimensions of Size, Position, Existence, and Count. Each test case consists of an image candidate set, which differs only in certain visual concepts, and a text candidate set, which differs only in the corresponding language concept.
git clone https://github.com/wjpoom/SPEC.git
cd SPEC/
pip install -e .
/path/to/save/data with a specified dir to store the data.import zipfile
import os
from huggingface_hub import hf_hub_download
data_root = '/path/to/save/data'
hf_hub_download(repo_id='wjpoom/SPEC', repo_type='dataset', filename='data.zip', local_dir=data_root)
with zipfile.ZipFile(os.path.join(data_root, 'data.zip'), 'r') as zip_ref:
zip_ref.extractall(os.path.join(data_root))
os.remove(os.path.join(data_root, 'data.zip'))
pip install open_clip_torch
mkdir checkpoints
huggingface-cli download wjpoom/SPEC-CLIP-ViT-B-32 --local-dir checkpoints/SPEC-CLIP-ViT-B-32
import torch
from PIL import Image
import open_clip
model, _, preprocess = open_clip.create_model_and_transforms('ViT-B-32', pretrained='checkpoints/SPEC-CLIP-ViT-B-32', load_weights_only=False)
model.eval()
tokenizer = open_clip.get_tokenizer('ViT-B-32')
image = preprocess(Image.open("assets/image.png")).unsqueeze(0)
text = tokenizer([
"the broccoli is situated above the backpack.",
"the broccoli is situated to the right of the backpack",
"the broccoli is positioned on the left of the backpack.",
"the broccoli is placed beneath the backpack."
])
with torch.no_grad(), torch.autocast("cuda"):
image_features = model.encode_image(image)
text_features = model.encode_text(text)
image_features /= image_features.norm(dim=-1, keepdim=True)
text_features /= text_features.norm(dim=-1, keepdim=True)
text_probs = (100.0 * image_features @ text_features.T).softmax(dim=-1)
print("Label probs:", text_probs)
Part of this repository is built upon ARO, thanks for the well-organized codebase.
Feel free to contact us if you have any questions or suggestions
Email (Wujian Peng): wjpeng24@m.fudan.edu.cn
If you use our code or data in this repo or find our work helpful, please consider giving a citation:
@inproceedings{peng2024synthesize,
title={Synthesize diagnose and optimize: Towards fine-grained vision-language understanding},
author={Peng, Wujian and Xie, Sicheng and You, Zuyao and Lan, Shiyi and Wu, Zuxuan},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={13279--13288},
year={2024}
}
Jupyter Notebook
98.6%
Python
1.4%