SHTUPLUS/Pix2Grp_CVPR2024

Python

71

6 commits

updated Nov 7, 2024

See the code

README

Official Implementation of "From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models"

Table of Contents

Introduction

Our paper "From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models" has been accepted by CVPR 2024.

Installation

  1. Creating conda environment and install pytorch
conda create -n pix2sgg python=3.8
conda activate pix2sgg

# CUDA 11.8
conda install pytorch==2.0.0 torchvision==0.15.0 pytorch-cuda=11.8 -c pytorch -c nvidia
# or CUDA 10.2
conda install pytorch==1.10.2 torchvision==0.11.3 cudatoolkit=10.2 -c pytorch
  1. Install other dependencies:
pip install -r requirements_pix2sgg.txt
# the hugging face version: v4.29.2

Our work is built upon LAVIS, sharing the majority of its requirements.

  1. Build Project
python setup.py build develop

Datasets

Check DATASET.md for instructions of dataset preprocessing.

Model Zoo

Open Vocabulary SGG

The model weight can be download from: https://huggingface.co/rj979797/PGSG-CVPR2024/tree/main

Novel+baseNovelcheckpoint
DatasetsmR50/100R50/100mR50/100
VG6.2/8.315.1/18.43.7/5.2vg_ov_sgg.pth
VG-SGCls9.7/13.826.8/33.25.1/7.7vg_ov_sgg.pth
PSG15.3/17.723.7/25.46.7/9.6psg_ov_sgg.pth

Close Vocabulary SGG

DatasetsmR50/100R50/100checkpoint
VG9.0/11.517.7/ 20.7vg_sgg.pth
PSG14.5/17.625.8/28.9psg_sgg.pth
VG-c10.4/12.720.3/23.6vg_sgg_close_clser.pth
PSG-c21.2/22.034.9/36.1psg_sgg_close_clser.pth

Training and Evaluation

Our PGSG is trained using the BLIP pre-trained weights, accessible here.

Ensure that the checkpoint path in the configuration file (*.yaml) is accurate before training or evaluation. During training, utilize the checkpoint specified by model.pretrained, while for evaluation, load the checkpoint specified by model.finetuned.

VG dataset

Open Vocabulary SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py  lavis/projects/blip/train/vrd_vg_ft_pgsg_ov.yaml --job-name VG-pgsg_ovsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_vg_pgsg_eval_ov.yaml --job-name VG-pgsg_stdsgg-eval 

Standard SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py  lavis/projects/blip/train/vrd_vg_ft_pgsg.yaml --job-name VG-pgsg_ovsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_vg_pgsg_eval.yaml --job-name VG-pgsg_stdsgg-eval 

PSG dataset

Open Vocabulary SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py --cfg-path lavis/projects/blip/train/vrd_psg_ft_pgsg_ov.yaml --job-name psg-pgsg_ovsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_psg_ov.yaml --job-name psg-pgsg_ovsgg-eval 

Standard SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py --cfg-path lavis/projects/blip/train/vrd_psg_ft_pgsg.yaml --job-name psg-pgsg_stdsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_psg_eval.yaml --job-name psg-pgsg_stdsgg-eval 

Paper and Citing

If you find this project helps your research, please kindly consider citing our papers in your publications.

@misc{li2024pixels,
    title={From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models},
    author={Rongjie Li and Songyang Zhang and Dahua Lin and Kai Chen and Xuming He},
    year={2024},
    eprint={2404.00906},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}

Acknowledge

This repository is built on LAVIS and borrows code from scene graph benchmarking framework from SGTR.

License

BSD 3-Clause License

SHTUPLUS/Pix2Grp_CVPR2024

Python

71

6 commits

updated Nov 7, 2024

See the code

README

Official Implementation of "From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models"

Table of Contents

Introduction

Our paper "From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models" has been accepted by CVPR 2024.

Installation

  1. Creating conda environment and install pytorch
conda create -n pix2sgg python=3.8
conda activate pix2sgg

# CUDA 11.8
conda install pytorch==2.0.0 torchvision==0.15.0 pytorch-cuda=11.8 -c pytorch -c nvidia
# or CUDA 10.2
conda install pytorch==1.10.2 torchvision==0.11.3 cudatoolkit=10.2 -c pytorch
  1. Install other dependencies:
pip install -r requirements_pix2sgg.txt
# the hugging face version: v4.29.2

Our work is built upon LAVIS, sharing the majority of its requirements.

  1. Build Project
python setup.py build develop

Datasets

Check DATASET.md for instructions of dataset preprocessing.

Model Zoo

Open Vocabulary SGG

The model weight can be download from: https://huggingface.co/rj979797/PGSG-CVPR2024/tree/main

Novel+baseNovelcheckpoint
DatasetsmR50/100R50/100mR50/100
VG6.2/8.315.1/18.43.7/5.2vg_ov_sgg.pth
VG-SGCls9.7/13.826.8/33.25.1/7.7vg_ov_sgg.pth
PSG15.3/17.723.7/25.46.7/9.6psg_ov_sgg.pth

Close Vocabulary SGG

DatasetsmR50/100R50/100checkpoint
VG9.0/11.517.7/ 20.7vg_sgg.pth
PSG14.5/17.625.8/28.9psg_sgg.pth
VG-c10.4/12.720.3/23.6vg_sgg_close_clser.pth
PSG-c21.2/22.034.9/36.1psg_sgg_close_clser.pth

Training and Evaluation

Our PGSG is trained using the BLIP pre-trained weights, accessible here.

Ensure that the checkpoint path in the configuration file (*.yaml) is accurate before training or evaluation. During training, utilize the checkpoint specified by model.pretrained, while for evaluation, load the checkpoint specified by model.finetuned.

VG dataset

Open Vocabulary SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py  lavis/projects/blip/train/vrd_vg_ft_pgsg_ov.yaml --job-name VG-pgsg_ovsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_vg_pgsg_eval_ov.yaml --job-name VG-pgsg_stdsgg-eval 

Standard SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py  lavis/projects/blip/train/vrd_vg_ft_pgsg.yaml --job-name VG-pgsg_ovsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_vg_pgsg_eval.yaml --job-name VG-pgsg_stdsgg-eval 

PSG dataset

Open Vocabulary SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py --cfg-path lavis/projects/blip/train/vrd_psg_ft_pgsg_ov.yaml --job-name psg-pgsg_ovsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_psg_ov.yaml --job-name psg-pgsg_ovsgg-eval 

Standard SGG

Training

python -m torch.distributed.run --master_port 13919 --nproc_per_node=4 train.py --cfg-path lavis/projects/blip/train/vrd_psg_ft_pgsg.yaml --job-name psg-pgsg_stdsgg

Evaluation

python -m torch.distributed.run --master_port 13958 --nproc_per_node=4 evaluate.py --cfg-path lavis/projects/blip/eval/rel_det_psg_eval.yaml --job-name psg-pgsg_stdsgg-eval 

Paper and Citing

If you find this project helps your research, please kindly consider citing our papers in your publications.

@misc{li2024pixels,
    title={From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models},
    author={Rongjie Li and Songyang Zhang and Dahua Lin and Kai Chen and Xuming He},
    year={2024},
    eprint={2404.00906},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}

Acknowledge

This repository is built on LAVIS and borrows code from scene graph benchmarking framework from SGTR.

License

BSD 3-Clause License

Languages

Python

77.6%

Jupyter Notebook

22.1%