Paper: https://arxiv.org/abs/2403.14599
Project Page: https://snap-research.github.io/MyVLM/
Code: https://github.com/snap-research/MyVLM
Example images for each object in our constructed dataset.
As part of our MyVLM code release, we have also released our object dataset introduced in the paper. This contains 29 user-specific objects, each containing ~10 images and 5 corresponding personalized captions for each image.
Your data should be organized using the following structure:
data_root
├── <concept_name>
│ ├── <image1>.jpg
│ ├── <image2>.jpg
│ ├── ...
│ ├── captions.json (or captions_augmented.json)
│ └── additional_llava_vqa_data.json (optional, used for personalized VQA using LLaVA, see next section).
└── <concept_name_2>
That is, the root directory should contain a sub-directory for each concept. Then, in each concept directory, you should have:
json file containing the captions for each image, named captions.json or captions_augmented.json.
This file should be in the following format:{
"<image1>.jpg": ["<caption1>", "<caption2>", ...],
"<image2>.jpg": ["<caption1>", "<caption2>", ...],
...
}
That is, we have a dictionary mapping each image path to a list of target captions. As described in the paper, at each optimization step we will randomly sample a caption from this list to use as the target caption for the image.
This sample code is made available by Snap Inc. for non-commercial, academic purposes only.
Please see the full license here.
6 commits
Paper: https://arxiv.org/abs/2403.14599
Project Page: https://snap-research.github.io/MyVLM/
Code: https://github.com/snap-research/MyVLM
Example images for each object in our constructed dataset.
As part of our MyVLM code release, we have also released our object dataset introduced in the paper. This contains 29 user-specific objects, each containing ~10 images and 5 corresponding personalized captions for each image.
Your data should be organized using the following structure:
data_root
├── <concept_name>
│ ├── <image1>.jpg
│ ├── <image2>.jpg
│ ├── ...
│ ├── captions.json (or captions_augmented.json)
│ └── additional_llava_vqa_data.json (optional, used for personalized VQA using LLaVA, see next section).
└── <concept_name_2>
That is, the root directory should contain a sub-directory for each concept. Then, in each concept directory, you should have:
json file containing the captions for each image, named captions.json or captions_augmented.json.
This file should be in the following format:{
"<image1>.jpg": ["<caption1>", "<caption2>", ...],
"<image2>.jpg": ["<caption1>", "<caption2>", ...],
...
}
That is, we have a dictionary mapping each image path to a list of target captions. As described in the paper, at each optimization step we will randomly sample a caption from this list to use as the target caption for the image.
This sample code is made available by Snap Inc. for non-commercial, academic purposes only.
Please see the full license here.
6 commits