PixMo-Count is a dataset of images paired with objects and their point locations in the image. It was built by running the Detic object detector on web images, and then filtering the data to improve accuracy and diversity. The val and test sets are human-verified and only contain counts from 2 to 10.
PixMo-Count is a part of the PixMo dataset collection and was used to augment the pointing capabilities of the Molmo family of models
Quick links:
data = datasets.load_dataset("allenai/pixmo-count", split="train")
Images are stored as URLs that will need to be downloaded separately. Note image URLs can be repeated in the data.
The points field contains the point x/y coordinates specified in pixels. Missing for the eval sets.
The label field contains the string name of the object being pointed at.
The count field contains the total count.
Image hashes are included to support double-checking that the downloaded image matches the annotated image. It can be checked like this:
from hashlib import sha256
import requests
example = data[0]
image_bytes = requests.get(example["image_url"]).content
byte_hash = sha256(image_bytes).hexdigest()
assert byte_hash == example["image_sha256"]
The test and val splits are human-verified but do not contain point information. We use them to evaluate counting capabilities of the Molmo models.
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
PixMo-Count is a dataset of images paired with objects and their point locations in the image. It was built by running the Detic object detector on web images, and then filtering the data to improve accuracy and diversity. The val and test sets are human-verified and only contain counts from 2 to 10.
PixMo-Count is a part of the PixMo dataset collection and was used to augment the pointing capabilities of the Molmo family of models
Quick links:
data = datasets.load_dataset("allenai/pixmo-count", split="train")
Images are stored as URLs that will need to be downloaded separately. Note image URLs can be repeated in the data.
The points field contains the point x/y coordinates specified in pixels. Missing for the eval sets.
The label field contains the string name of the object being pointed at.
The count field contains the total count.
Image hashes are included to support double-checking that the downloaded image matches the annotated image. It can be checked like this:
from hashlib import sha256
import requests
example = data[0]
image_bytes = requests.get(example["image_url"]).content
byte_hash = sha256(image_bytes).hexdigest()
assert byte_hash == example["image_sha256"]
The test and val splits are human-verified but do not contain point information. We use them to evaluate counting capabilities of the Molmo models.
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.