· huggingface · github · license
SpatialMQA is a manually annotated dataset designed for multimodal spatial relation reasoning in a multiple-choice question & answer format. The dataset includes 5,392 samples collected from COCO2017, covering 128 subject and object types, without bounding boxes. To address the limitations of existing datasets, we clearly define annotation guidelines for SpatialMQA, including standardizing the objective world as the coordinate system and avoiding questions that can be answered solely by the question itself.
The following figures list some classic examples in our dataset. You can click out Examples:1~4/ and Examples:5~8/ to view the details.
The following table Splits/ lists the detailed information statistics of the splited dataset.
You can find our dataset through the following path (Dataset/dataset) for more details.
Due to the fact that only redirecting to the specified file is valid in anonymous links, redirecting to the specified directory is invalid. Therefore, we use bold and italicized font to indicate the markings of all specified directories, making it easier for reviewers to search. Thank you!
Our dataset has been officially released on the Hugging Face. It is available at https://huggingface.co/datasets/liuziyan/SpatialMQA.
Alternatively, you can download it from GitHub by following the steps below:
We use a subset of COCO-2017's images. The following script download COCO-2017's test sets images then put them into a single fodler Dataset/COCO2017/.
cd Dataset/
wget http://images.cocodataset.org/zips/test2017.zip
unzip test2017.zip
mv test2017 COCO2017 && rm -r test2017
Copy only relevant images to relevant_images/.
mkdir relevant_images
cd tool
python select_revlevant_images.py
Alternatively, you could also browse individual images online directly using the key "image" in single json data.
(Through COCO's open source link, 'http://images.cocodataset.org/test2017/' + 'image_name'. For example: http://images.cocodataset.org/test2017/000000195921.jpg.)
As reported in the folloeing table, SpatialMQA contains 5,392 samples, divided into training, validation, and test sets according to a 7:1:2 ratio.
All the splited data sets are in the directory (Dataset/dataset).
In addition, we have selected the part of the data set that contains invisible subjects or objects and placed it in the file invisible/ to facilitate readers' research.
Each jsonl file is of the following format:
{"image": "000000000933.jpg", "question": "Where is the fork located relative to the pizza?", "options": ["on/above", "below", "in front of", "behind", "left of", "right of"], "answer": "right of"}
{"image": "000000100633.jpg", "question": "If you are the cyclist in the image, where is the dog located relative to you?", "options": ["in front of", "behind", "left of", "right of"], "answer": "behind"}
{"image": "000000070986.jpg", "question": "If you are the driver of the bus in the image, from your perspective, where is the red car located relative to the bus?", "options": ["in front of", "behind", "left of", "right of"], "answer": "left of"}
{"..."}
Each line is an individual data point.
image denotes name of the image in COCO. question is the question with manual annotation, options is reasonable combinations of six spatial relationships:(on/above, below, in front of, behind, left of, right of. answer is the annotation based on the objective world.
Our dataset is expanded based on the categories included in the COCO dataset. There are 113 subject types and one additional type for subjects with five or fewer samples in our dataset, and 84 object types and one additional type for objects with five or fewer samples. Due to the overlap between subject and object types, we have a total of 128 distinct subject and object types. You can see all of them in the file S & O types/.
We have disclosed the inference code for the model in the directory (Code/experiment), as well as the fine-tuning code in the directory (Code/finetune).
nohup python blip-vqa-base.py > log/blip_exp.log 2>1& &
nohup python blip-vqa-base_finetuned.py > log/blip_finetuned_exp.log 2>1& &
nohup python blip2-opt-2.7b.py > log/blip2_exp.log 2>1& &
nohup python blip2-lora.py > log/blip2_lora_exp.log 2>1& &
nohup python instructblip-flan-t5-xl.py > log/instructblip_exp.log 2>1& &
nohup python instructblip-lora.py > log/instructblip_lora_exp.log 2>1& &
nohup python idefics_new.py > log/idefics_exp.log 2>1& &
nohup python idefics_lora.py > log/idefics_lora_exp.log 2>1& &
nohup python spatial_test_llava.py > log/llava_exp.log 2>1& &
nohup python spatial_test_llava_lora.py > log/llava_lora_exp.log 2>1& &
nohup python spatial_test_mplug.py > log/mplug_exp.log 2>1& &
nohup python spatial_test_mplug_lora.py > log/mplug_lora_exp.log 2>1& &
nohup python spacellava_test.py > log/spacellava_exp.log 2>1& &
nohup python spacellava_lora_test.py > log/spacellava_lora_exp.log 2>1& &
Due to the large amount of open source model code, you need to download it yourself through channels or call it directly from platforms such as huggingface.
nohup python blip-vqa-base.py > log/blip_train.log 2>1& &
nohup python blip2-lora.py > log/blip2_train.log 2>1& &
nohup python instructblip-lora.py > log/instructblip_train.log 2>1& &
nohup python idefics.py > log/idefics_train.log 2>1& &
nohup bash llava_lora_train.sh > log/llava_train.log 2>1& &
nohup bash spacellava_lora_train.sh > log/spacellava_train.log 2>1& &
nohup bash mPLUG_Owl_train_it.sh > log/mplug_train.log 2>1& &
python gemini_text_only.py
python gemini_zero_shot.py
python gemini_1_shot.py
python gemini_2_shot.py
python gemini_3_shot.py
python gpt4_text_only.py
python gpt4_zero_shot.py
python gpt4_1_shot.py
python gpt4_2_shot.py
python gpt4_3_shot.py
Gemini needs to apply on the official website, and GPT4 needs to be purchased on the official website.
You can process the results of model inference through the code we provide to calculate overall accuracy, overall P, R, F1 indicators, accuracy based on relationship categories, and accuracy based on rules. We integrate the calculation process into the Python files in the directory (Code/eval):
python calculate_prf1.py
python calculate_xyz.py
python calculate_result_rule.py
The environment configuration required for debugging code is placed in directory (Code/requirement)
The requirements of models blip, blip2 and instructblip, are all in the file requirement_blip.txt/
This project is licensed under the Apache-2.0 License.
696 followers · starred Aug 2025
Python
97.5%
Shell
2.5%
· huggingface · github · license
SpatialMQA is a manually annotated dataset designed for multimodal spatial relation reasoning in a multiple-choice question & answer format. The dataset includes 5,392 samples collected from COCO2017, covering 128 subject and object types, without bounding boxes. To address the limitations of existing datasets, we clearly define annotation guidelines for SpatialMQA, including standardizing the objective world as the coordinate system and avoiding questions that can be answered solely by the question itself.
The following figures list some classic examples in our dataset. You can click out Examples:1~4/ and Examples:5~8/ to view the details.
The following table Splits/ lists the detailed information statistics of the splited dataset.
You can find our dataset through the following path (Dataset/dataset) for more details.
Due to the fact that only redirecting to the specified file is valid in anonymous links, redirecting to the specified directory is invalid. Therefore, we use bold and italicized font to indicate the markings of all specified directories, making it easier for reviewers to search. Thank you!
Our dataset has been officially released on the Hugging Face. It is available at https://huggingface.co/datasets/liuziyan/SpatialMQA.
Alternatively, you can download it from GitHub by following the steps below:
We use a subset of COCO-2017's images. The following script download COCO-2017's test sets images then put them into a single fodler Dataset/COCO2017/.
cd Dataset/
wget http://images.cocodataset.org/zips/test2017.zip
unzip test2017.zip
mv test2017 COCO2017 && rm -r test2017
Copy only relevant images to relevant_images/.
mkdir relevant_images
cd tool
python select_revlevant_images.py
Alternatively, you could also browse individual images online directly using the key "image" in single json data.
(Through COCO's open source link, 'http://images.cocodataset.org/test2017/' + 'image_name'. For example: http://images.cocodataset.org/test2017/000000195921.jpg.)
As reported in the folloeing table, SpatialMQA contains 5,392 samples, divided into training, validation, and test sets according to a 7:1:2 ratio.
All the splited data sets are in the directory (Dataset/dataset).
In addition, we have selected the part of the data set that contains invisible subjects or objects and placed it in the file invisible/ to facilitate readers' research.
Each jsonl file is of the following format:
{"image": "000000000933.jpg", "question": "Where is the fork located relative to the pizza?", "options": ["on/above", "below", "in front of", "behind", "left of", "right of"], "answer": "right of"}
{"image": "000000100633.jpg", "question": "If you are the cyclist in the image, where is the dog located relative to you?", "options": ["in front of", "behind", "left of", "right of"], "answer": "behind"}
{"image": "000000070986.jpg", "question": "If you are the driver of the bus in the image, from your perspective, where is the red car located relative to the bus?", "options": ["in front of", "behind", "left of", "right of"], "answer": "left of"}
{"..."}
Each line is an individual data point.
image denotes name of the image in COCO. question is the question with manual annotation, options is reasonable combinations of six spatial relationships:(on/above, below, in front of, behind, left of, right of. answer is the annotation based on the objective world.
Our dataset is expanded based on the categories included in the COCO dataset. There are 113 subject types and one additional type for subjects with five or fewer samples in our dataset, and 84 object types and one additional type for objects with five or fewer samples. Due to the overlap between subject and object types, we have a total of 128 distinct subject and object types. You can see all of them in the file S & O types/.
We have disclosed the inference code for the model in the directory (Code/experiment), as well as the fine-tuning code in the directory (Code/finetune).
nohup python blip-vqa-base.py > log/blip_exp.log 2>1& &
nohup python blip-vqa-base_finetuned.py > log/blip_finetuned_exp.log 2>1& &
nohup python blip2-opt-2.7b.py > log/blip2_exp.log 2>1& &
nohup python blip2-lora.py > log/blip2_lora_exp.log 2>1& &
nohup python instructblip-flan-t5-xl.py > log/instructblip_exp.log 2>1& &
nohup python instructblip-lora.py > log/instructblip_lora_exp.log 2>1& &
nohup python idefics_new.py > log/idefics_exp.log 2>1& &
nohup python idefics_lora.py > log/idefics_lora_exp.log 2>1& &
nohup python spatial_test_llava.py > log/llava_exp.log 2>1& &
nohup python spatial_test_llava_lora.py > log/llava_lora_exp.log 2>1& &
nohup python spatial_test_mplug.py > log/mplug_exp.log 2>1& &
nohup python spatial_test_mplug_lora.py > log/mplug_lora_exp.log 2>1& &
nohup python spacellava_test.py > log/spacellava_exp.log 2>1& &
nohup python spacellava_lora_test.py > log/spacellava_lora_exp.log 2>1& &
Due to the large amount of open source model code, you need to download it yourself through channels or call it directly from platforms such as huggingface.
nohup python blip-vqa-base.py > log/blip_train.log 2>1& &
nohup python blip2-lora.py > log/blip2_train.log 2>1& &
nohup python instructblip-lora.py > log/instructblip_train.log 2>1& &
nohup python idefics.py > log/idefics_train.log 2>1& &
nohup bash llava_lora_train.sh > log/llava_train.log 2>1& &
nohup bash spacellava_lora_train.sh > log/spacellava_train.log 2>1& &
nohup bash mPLUG_Owl_train_it.sh > log/mplug_train.log 2>1& &
python gemini_text_only.py
python gemini_zero_shot.py
python gemini_1_shot.py
python gemini_2_shot.py
python gemini_3_shot.py
python gpt4_text_only.py
python gpt4_zero_shot.py
python gpt4_1_shot.py
python gpt4_2_shot.py
python gpt4_3_shot.py
Gemini needs to apply on the official website, and GPT4 needs to be purchased on the official website.
You can process the results of model inference through the code we provide to calculate overall accuracy, overall P, R, F1 indicators, accuracy based on relationship categories, and accuracy based on rules. We integrate the calculation process into the Python files in the directory (Code/eval):
python calculate_prf1.py
python calculate_xyz.py
python calculate_result_rule.py
The environment configuration required for debugging code is placed in directory (Code/requirement)
The requirements of models blip, blip2 and instructblip, are all in the file requirement_blip.txt/
This project is licensed under the Apache-2.0 License.
696 followers · starred Aug 2025
Python
97.5%
Shell
2.5%