csebuetnlp/IllusionVQA

This repository contains the data and code of the paper titled "IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models"

Jupyter Notebook

24

75 commits

updated Apr 27, 2025

See the code

README

IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

Code for the Paper "IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models".

For more details, please refer to the project page: https://illusionvqa.github.io.

πŸ”” If you have any questions or suggestions, please don't hesitate to let us know. You can post an issue on this repository or mail us directly.

Project Page | Paper | πŸ€— IllusionVQA-Comprehension | πŸ€— IllusionVQA-Soft-Localization

πŸ‘€ TL;DR

IllusionVQA is a dataset of optical illusions and hard-to-interpret scenes designed to test the capability of Vision Language Models in comprehension and soft localization tasks. GPT4V achieved 62.99% accuracy on comprehension and 49.7% on localization, while humans achieved 91.03% and 100% respectively.

πŸ’₯ News πŸ’₯

  • [2024.08.31] πŸ’₯ Gemini-1.5-Pro sets new SOTA on both Comprehension and Soft-Localization! With a significant lead in Comprehension (71% vs. second place 67%). Gemini-1.5-Flash places 7th and 4th respectively.
  • [2024.08.16] πŸ’₯ Claude 3.5 Sonnet achieves 2nd place on comprehension with 66.44! Learn more at the Anthropic blog.
  • [2024.08.16] πŸ’₯ OpenAI's GPT-4o achieves new SOTA on IllusionVQA with 67.12% on Comprehension and 53.3% on Soft Localization! Learn more at the OpenAI blog.
  • [2024.07.28] πŸš€ InternVL2 achieves 45.06% on Comprehension and 28.3% on Soft Localization, scoring the best among open source models. πŸŽ‰ Congratulations!
  • [2024.07.09] 🌟 Our IllusionVQA paper has been accepted at COLM 2024 (acceptance rate 28.8%)! πŸŽ‰ Cheers!
  • [2024.05.28] ✨ Our work was featured by Scientific American. Thanks! ✨
  • [2024.03.28] πŸš€ Our project page is live at https://illusionvqa.github.io.
  • [2024.03.27] Our dataset is now accessible at Papers With Code.
  • [2024.03.26] Our dataset is now accessible at Huggingface Datasets! 🧠 Comprehension and πŸ”Ž Soft Localization.
  • [2024.03.26] Our paper is now accessible at https://arxiv.org/abs/2403.15952.

πŸ† Results πŸ†

For the latest results, checkout out the leaderboard in the project page.

IllusionVQA-Comprehension

Class#0-shot4-shotHuman
I-BLIPLLaVACogGeminiGPT4VGeminiGPT4V
Impossible Object13434.2243.2844.0356.7255.2256.7258.9698.51
Real-Scene6426.5642.1934.3846.8857.8146.8854.6998.44
Size4626.0919.5713.0445.6558.7052.1769.5763.04
Hidden4544.4442.2242.2242.2251.1148.8946.67100
Deceptive Design3737.8443.2445.9564.8670.2767.5672.9794.59
Angle Illusion2630.7738.4630.7753.8569.235084.6284.62
Color2330.4326.0930.4317.3969.5717.3982.6160.87
Edited-Scene2142.8661.9042.8666.6771.4366.6780.95100
Upside-Down742.8671.4371.4357.1471.4357.1471.43100
Pos.-Neg. Space757.4142.8671.4385.7157.1471.4385.71100
Circle-Spiral633.330.0016.6733.335033.3333.3366.67
Miscellaneous1936.8442.1142.1152.6342.1157.8942.1189.47
Total43534.254038.1651.2658.8552.8762.9991.03

New Results [13 July 2024]

Class#0-shot4-shotHuman
gpt4ogpt4o
Impossible Object13463.4361.9498.51
Real-Scene6464.0657.8198.44
Size4645.6593.4763.04
Hidden4566.6748.89100
Deceptive Design3772.9778.3894.59
Angle Illusion2650.0080.7784.62
Color2352.1778.2660.87
Edited-Scene2180.9585.71100
Upside-Down771.4342.86100
Pos.-Neg. Space785.7171.43100
Circle-Spiral650.0050.0066.67
Miscellaneous1952.6352.6389.47
Total43562.5367.1291.03

IllusonvQA-Soft-Localization

VLMPrompt TypeAccuracy
InstructBLIP0-shot24.3
LLaVA-1.50-shot24.8
CogVLM0-shot28
GPT4V0-shot40
4-shot46
4-shot + CoT49.7
Gemini Pro0-shot43.5
4-shot41.8
4-shot + CoT33.9
Human100

πŸ“– Usage

from datasets import load_dataset
import base64
from openai import OpenAI
import os
os.environ["OPENAI_API_KEY"] = "YOUR_API_KEY"

def encode_image(pil_image):
    temp_name = "temp.jpg"
    pil_image.save(temp_name)
    with open(temp_name, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

def construct_mcq(options, correct_option):
    correct_option_letter = None
    i = "a"
    mcq = ""
    for option in options:
        if option == correct_option:
            correct_option_letter = i
        mcq += f"{i}. {option}\n"
        i = chr(ord(i) + 1)
    mcq = mcq[:-1]
    return mcq, correct_option_letter

def add_row(content, data, i, with_answer=False):  
    mcq, correct_option_letter = construct_mcq(data["options"], data["answer"])
    content.append({ "type": "text",
            "text": "Image " + str(i) + ": " + data["question"] + "\n" + mcq })
    content.append({ "type": "image_url",
            "image_url": {"url": f"data:image/jpeg;base64,{encode_image(data['image'])}",
                "detail": "low"}})
    if with_answer:
        content.append({"type": "text", "text": "Answer {}: ".format(i) + correct_option_letter})
    else:
        content.append({"type": "text", "text": "Answer {}: ".format(i), })
    return content

dataset = load_dataset("csebuetnlp/illusionVQA-Comprehension")
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

content = [{
        "type": "text",
        "text": "You'll be given an image, an instruction and some choices. You have to select the correct one. Do not explain your reasoning. Answer with the option's letter from the given choices directly. Here are a few examples:",
    }]

### Add a few examples
for i, data in enumerate(dataset["train"], 1):
    content = add_row(content, data, i, with_answer=True)

content.append({"type": "text", "text": "Now you try it!",})

next_idx = i + 1

### Add the test data
test_data = dataset["test"][0]
content_t = add_row(content.copy(), test_data, next_idx, with_answer=False)

### Get the answer from GPT-4
response = client.chat.completions.create(
    model="gpt-4-vision-preview",
    messages=[{"role": "user","content": content_t,}],
    max_tokens=5,
)
gpt4_answer = response.choices[0].message.content
print(gpt4_answer)

πŸ“œ License

This dataset is made available for non-commercial research purposes only under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). The dataset may not be used for training models. The dataset contains images collected from the internet. While permission has been obtained from some of the images' creators, permission has not yet been received from all creators. If you believe any image in this dataset is used without proper permission and you are the copyright holder, please email Haz Sameen Shahgir to request the removal of the image from the dataset.

The dataset creator makes no representations or warranties regarding the copyright status of the images in the dataset. The dataset creator shall not be held liable for any unauthorized use of copyrighted material that may be contained in the dataset.

You agree to the terms and conditions specified in this license by downloading or using this dataset. If you do not agree with these terms, do not download or use the dataset.

Creative Commons License

βœ… Cite

@inproceedings{
shahgir2024illusionvqa,
title={Illusion{VQA}: A Challenging Optical Illusion Dataset for Vision Language Models},
author={Haz Sameen Shahgir and Khondker Salman Sayeed and Abhik Bhattacharjee and Wasi Uddin Ahmad and Yue Dong and Rifat Shahriyar},
booktitle={First Conference on Language Modeling},
year={2024},
url={https://openreview.net/forum?id=7ysaJGs7zY}
}
optical-illusions
visual-language-models
vqa
vqa-dataset

csebuetnlp/IllusionVQA

This repository contains the data and code of the paper titled "IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models"

Jupyter Notebook

24

75 commits

updated Apr 27, 2025

See the code

README

IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

Code for the Paper "IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models".

For more details, please refer to the project page: https://illusionvqa.github.io.

πŸ”” If you have any questions or suggestions, please don't hesitate to let us know. You can post an issue on this repository or mail us directly.

Project Page | Paper | πŸ€— IllusionVQA-Comprehension | πŸ€— IllusionVQA-Soft-Localization

πŸ‘€ TL;DR

IllusionVQA is a dataset of optical illusions and hard-to-interpret scenes designed to test the capability of Vision Language Models in comprehension and soft localization tasks. GPT4V achieved 62.99% accuracy on comprehension and 49.7% on localization, while humans achieved 91.03% and 100% respectively.

πŸ’₯ News πŸ’₯

  • [2024.08.31] πŸ’₯ Gemini-1.5-Pro sets new SOTA on both Comprehension and Soft-Localization! With a significant lead in Comprehension (71% vs. second place 67%). Gemini-1.5-Flash places 7th and 4th respectively.
  • [2024.08.16] πŸ’₯ Claude 3.5 Sonnet achieves 2nd place on comprehension with 66.44! Learn more at the Anthropic blog.
  • [2024.08.16] πŸ’₯ OpenAI's GPT-4o achieves new SOTA on IllusionVQA with 67.12% on Comprehension and 53.3% on Soft Localization! Learn more at the OpenAI blog.
  • [2024.07.28] πŸš€ InternVL2 achieves 45.06% on Comprehension and 28.3% on Soft Localization, scoring the best among open source models. πŸŽ‰ Congratulations!
  • [2024.07.09] 🌟 Our IllusionVQA paper has been accepted at COLM 2024 (acceptance rate 28.8%)! πŸŽ‰ Cheers!
  • [2024.05.28] ✨ Our work was featured by Scientific American. Thanks! ✨
  • [2024.03.28] πŸš€ Our project page is live at https://illusionvqa.github.io.
  • [2024.03.27] Our dataset is now accessible at Papers With Code.
  • [2024.03.26] Our dataset is now accessible at Huggingface Datasets! 🧠 Comprehension and πŸ”Ž Soft Localization.
  • [2024.03.26] Our paper is now accessible at https://arxiv.org/abs/2403.15952.

πŸ† Results πŸ†

For the latest results, checkout out the leaderboard in the project page.

IllusionVQA-Comprehension

Class#0-shot4-shotHuman
I-BLIPLLaVACogGeminiGPT4VGeminiGPT4V
Impossible Object13434.2243.2844.0356.7255.2256.7258.9698.51
Real-Scene6426.5642.1934.3846.8857.8146.8854.6998.44
Size4626.0919.5713.0445.6558.7052.1769.5763.04
Hidden4544.4442.2242.2242.2251.1148.8946.67100
Deceptive Design3737.8443.2445.9564.8670.2767.5672.9794.59
Angle Illusion2630.7738.4630.7753.8569.235084.6284.62
Color2330.4326.0930.4317.3969.5717.3982.6160.87
Edited-Scene2142.8661.9042.8666.6771.4366.6780.95100
Upside-Down742.8671.4371.4357.1471.4357.1471.43100
Pos.-Neg. Space757.4142.8671.4385.7157.1471.4385.71100
Circle-Spiral633.330.0016.6733.335033.3333.3366.67
Miscellaneous1936.8442.1142.1152.6342.1157.8942.1189.47
Total43534.254038.1651.2658.8552.8762.9991.03

New Results [13 July 2024]

Class#0-shot4-shotHuman
gpt4ogpt4o
Impossible Object13463.4361.9498.51
Real-Scene6464.0657.8198.44
Size4645.6593.4763.04
Hidden4566.6748.89100
Deceptive Design3772.9778.3894.59
Angle Illusion2650.0080.7784.62
Color2352.1778.2660.87
Edited-Scene2180.9585.71100
Upside-Down771.4342.86100
Pos.-Neg. Space785.7171.43100
Circle-Spiral650.0050.0066.67
Miscellaneous1952.6352.6389.47
Total43562.5367.1291.03

IllusonvQA-Soft-Localization

VLMPrompt TypeAccuracy
InstructBLIP0-shot24.3
LLaVA-1.50-shot24.8
CogVLM0-shot28
GPT4V0-shot40
4-shot46
4-shot + CoT49.7
Gemini Pro0-shot43.5
4-shot41.8
4-shot + CoT33.9
Human100

πŸ“– Usage

from datasets import load_dataset
import base64
from openai import OpenAI
import os
os.environ["OPENAI_API_KEY"] = "YOUR_API_KEY"

def encode_image(pil_image):
    temp_name = "temp.jpg"
    pil_image.save(temp_name)
    with open(temp_name, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

def construct_mcq(options, correct_option):
    correct_option_letter = None
    i = "a"
    mcq = ""
    for option in options:
        if option == correct_option:
            correct_option_letter = i
        mcq += f"{i}. {option}\n"
        i = chr(ord(i) + 1)
    mcq = mcq[:-1]
    return mcq, correct_option_letter

def add_row(content, data, i, with_answer=False):  
    mcq, correct_option_letter = construct_mcq(data["options"], data["answer"])
    content.append({ "type": "text",
            "text": "Image " + str(i) + ": " + data["question"] + "\n" + mcq })
    content.append({ "type": "image_url",
            "image_url": {"url": f"data:image/jpeg;base64,{encode_image(data['image'])}",
                "detail": "low"}})
    if with_answer:
        content.append({"type": "text", "text": "Answer {}: ".format(i) + correct_option_letter})
    else:
        content.append({"type": "text", "text": "Answer {}: ".format(i), })
    return content

dataset = load_dataset("csebuetnlp/illusionVQA-Comprehension")
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

content = [{
        "type": "text",
        "text": "You'll be given an image, an instruction and some choices. You have to select the correct one. Do not explain your reasoning. Answer with the option's letter from the given choices directly. Here are a few examples:",
    }]

### Add a few examples
for i, data in enumerate(dataset["train"], 1):
    content = add_row(content, data, i, with_answer=True)

content.append({"type": "text", "text": "Now you try it!",})

next_idx = i + 1

### Add the test data
test_data = dataset["test"][0]
content_t = add_row(content.copy(), test_data, next_idx, with_answer=False)

### Get the answer from GPT-4
response = client.chat.completions.create(
    model="gpt-4-vision-preview",
    messages=[{"role": "user","content": content_t,}],
    max_tokens=5,
)
gpt4_answer = response.choices[0].message.content
print(gpt4_answer)

πŸ“œ License

This dataset is made available for non-commercial research purposes only under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). The dataset may not be used for training models. The dataset contains images collected from the internet. While permission has been obtained from some of the images' creators, permission has not yet been received from all creators. If you believe any image in this dataset is used without proper permission and you are the copyright holder, please email Haz Sameen Shahgir to request the removal of the image from the dataset.

The dataset creator makes no representations or warranties regarding the copyright status of the images in the dataset. The dataset creator shall not be held liable for any unauthorized use of copyrighted material that may be contained in the dataset.

You agree to the terms and conditions specified in this license by downloading or using this dataset. If you do not agree with these terms, do not download or use the dataset.

Creative Commons License

βœ… Cite

@inproceedings{
shahgir2024illusionvqa,
title={Illusion{VQA}: A Challenging Optical Illusion Dataset for Vision Language Models},
author={Haz Sameen Shahgir and Khondker Salman Sayeed and Abhik Bhattacharjee and Wasi Uddin Ahmad and Yue Dong and Rifat Shahriyar},
booktitle={First Conference on Language Modeling},
year={2024},
url={https://openreview.net/forum?id=7ysaJGs7zY}
}
optical-illusions
visual-language-models
vqa
vqa-dataset