HarborYuan/VRT-Eval

Dataset

0

stars

6

commits

3

linked in READMEs

May 7, 2026

updated

reasoning-segmentation
referring-expression-segmentation
sa2va
visual-reasoning
vrt

README

VRT-Eval

Human-verified evaluation benchmark for the Visual Reasoning Tracer (VRT) task — given an image and a natural-language question, a model must (1) trace the supporting reasoning objects and (2) point at the answer object(s) with segmentation masks.

Built on Sa2VA and tfd-utils. The benchmark is released alongside the Sa2VA / VRT codebase, and the data is packaged as a single random-access TFRecord via tfd-utils. Both are required to load and evaluate.

The training counterpart lives at HarborYuan/VisualReasoningTracer.

Quick stats

Samples304
Filevrt_eval.tfrecord (276 MB) + vrt_eval.index
Source imagesSA-1B
Question categories#visf (visual feature), #loc (localization), #comp (composition/relation), #func (function/affordance)
Avg. candidate objects per image29.5 (2–138)
Avg. ground-truth answer objects1.26 (1–6)
Avg. ground-truth reasoning objects3.34 (1–10)
Annotator confidence1–5 (87% are 4 or 5)

Each sample carries one or more category tags from {#visf, #loc, #comp, #func}; categories overlap, so the per-tag counts (126 / 132 / 102 / 128) sum above 304.

Schema

The dataset is packaged as a single TFRecord with a sidecar .index for random access via tfd-utils. Each example has the following fields:

FeatureTypeDescription
keybytesUnique sample id
imagebytesJPEG-encoded RGB image
questionbytes (utf-8)Natural-language question
objects_info_jsonbytes (utf-8 JSON)Dict of {obj_id: {mask: COCO-RLE, original_id, caption}} — the candidate object pool
human_labeled_r_objsint64 listIndices into objects_info for reasoning objects
human_labeled_a_objsint64 listIndices into objects_info for answer objects
human_confidenceint64Annotator confidence, 1–5
class_idsbytes listQuestion category tag(s), e.g. ["#visf", "#loc"]
human_reasoningbytes (utf-8)Optional human-written reasoning trace (may be empty)
human_answer_captionbytes (utf-8)Optional human-written answer caption (may be empty)

Masks are stored as COCO RLE; decode with pycocotools.mask.decode.

Download

huggingface-cli download HarborYuan/VRT-Eval \
  --repo-type dataset \
  --local-dir-use-symlinks False \
  --local-dir ./data/VRT-Eval

Load

A reference random-access loader ships in the Sa2VA repo at projects/vrt_sa2va/evaluation/packed_vrt_eval_dataloader.py:

from projects.vrt_sa2va.evaluation.packed_vrt_eval_dataloader import PackedVRTEvalDataset

ds = PackedVRTEvalDataset("data/VRT-Eval/vrt_eval.tfrecord")
for key in ds.get_all_keys():
    sample = ds.get_sample(key)
    # sample = {key, image (PIL.Image), question, objects_info,
    #           human_labeled_r_objs, human_labeled_a_objs,
    #           human_confidence, class_ids,
    #           human_reasoning, human_answer_caption}

objects_info[obj_id]['mask'] is a decoded H×W np.uint8 binary mask.

Evaluate

End-to-end on a Sa2VA / VRT checkpoint already in HuggingFace format:

PYTHONPATH=. uv run --extra latest \
  projects/vrt_sa2va/evaluation/vrt_eval_single.py \
  <hf_model_dir> --gpus 8

Or fold conversion + eval into one step from a training .pth:

uv run --extra latest bash ./vrt_fast_eval.sh \
  work_dirs/<exp>/iter_XXXX.pth 8

Predicted masks in the model output are split into <think> (reasoning) and <answer> segments and matched against human_labeled_r_objs / human_labeled_a_objs via Hungarian assignment on IoU.

Evaluating closed-source LLMs (Gemini / GPT / OpenRouter-hosted Qwen)

For models without open weights, we provide an OpenRouter + SAM2 evaluator:

📁 projects/vrt_sa2va/eval_openrouter/ — VER mask evaluation via OpenRouter

It (1) prompts the LLM through OpenRouter for bounding boxes, (2) feeds the boxes into SAM2 to produce masks, then (3) scores those masks against the human-labelled GT in this dataset using the same metrics as vrt_eval_single.py.

Metrics

MetricMeaning
answer_miouMain metric. Mean IoU on the answer mask(s)
reasoning_miouMean IoU on the reasoning mask(s)
LQLogical Quality — % of predictions with IoU > 0.5
SQSegmentation Quality — mean IoU among predictions with IoU > 0.5

Question prompt

Match the training prompt for fair comparison:

<image>
{question}

You should first think about the reasoning process in the mind and then
provides the user with the answer. Please respond with segmentation mask in
both the thinking process and the answer.

The reasoning process and answer are enclosed within <think> </think> and
<answer> </answer> tags, respectively, i.e., <think> reasoning process here
</think> <answer> answers here </answer>.

License

Annotations are released under CC BY-NC 4.0. Source images come from SA-1B and are subject to the SA-1B license.

Contributors

HarborYuan

6 commits

HarborYuan/VRT-Eval

Dataset

0

stars

6

commits

3

linked in READMEs

May 7, 2026

updated

reasoning-segmentation
referring-expression-segmentation
sa2va
visual-reasoning
vrt

README

VRT-Eval

Human-verified evaluation benchmark for the Visual Reasoning Tracer (VRT) task — given an image and a natural-language question, a model must (1) trace the supporting reasoning objects and (2) point at the answer object(s) with segmentation masks.

Built on Sa2VA and tfd-utils. The benchmark is released alongside the Sa2VA / VRT codebase, and the data is packaged as a single random-access TFRecord via tfd-utils. Both are required to load and evaluate.

The training counterpart lives at HarborYuan/VisualReasoningTracer.

Quick stats

Samples304
Filevrt_eval.tfrecord (276 MB) + vrt_eval.index
Source imagesSA-1B
Question categories#visf (visual feature), #loc (localization), #comp (composition/relation), #func (function/affordance)
Avg. candidate objects per image29.5 (2–138)
Avg. ground-truth answer objects1.26 (1–6)
Avg. ground-truth reasoning objects3.34 (1–10)
Annotator confidence1–5 (87% are 4 or 5)

Each sample carries one or more category tags from {#visf, #loc, #comp, #func}; categories overlap, so the per-tag counts (126 / 132 / 102 / 128) sum above 304.

Schema

The dataset is packaged as a single TFRecord with a sidecar .index for random access via tfd-utils. Each example has the following fields:

FeatureTypeDescription
keybytesUnique sample id
imagebytesJPEG-encoded RGB image
questionbytes (utf-8)Natural-language question
objects_info_jsonbytes (utf-8 JSON)Dict of {obj_id: {mask: COCO-RLE, original_id, caption}} — the candidate object pool
human_labeled_r_objsint64 listIndices into objects_info for reasoning objects
human_labeled_a_objsint64 listIndices into objects_info for answer objects
human_confidenceint64Annotator confidence, 1–5
class_idsbytes listQuestion category tag(s), e.g. ["#visf", "#loc"]
human_reasoningbytes (utf-8)Optional human-written reasoning trace (may be empty)
human_answer_captionbytes (utf-8)Optional human-written answer caption (may be empty)

Masks are stored as COCO RLE; decode with pycocotools.mask.decode.

Download

huggingface-cli download HarborYuan/VRT-Eval \
  --repo-type dataset \
  --local-dir-use-symlinks False \
  --local-dir ./data/VRT-Eval

Load

A reference random-access loader ships in the Sa2VA repo at projects/vrt_sa2va/evaluation/packed_vrt_eval_dataloader.py:

from projects.vrt_sa2va.evaluation.packed_vrt_eval_dataloader import PackedVRTEvalDataset

ds = PackedVRTEvalDataset("data/VRT-Eval/vrt_eval.tfrecord")
for key in ds.get_all_keys():
    sample = ds.get_sample(key)
    # sample = {key, image (PIL.Image), question, objects_info,
    #           human_labeled_r_objs, human_labeled_a_objs,
    #           human_confidence, class_ids,
    #           human_reasoning, human_answer_caption}

objects_info[obj_id]['mask'] is a decoded H×W np.uint8 binary mask.

Evaluate

End-to-end on a Sa2VA / VRT checkpoint already in HuggingFace format:

PYTHONPATH=. uv run --extra latest \
  projects/vrt_sa2va/evaluation/vrt_eval_single.py \
  <hf_model_dir> --gpus 8

Or fold conversion + eval into one step from a training .pth:

uv run --extra latest bash ./vrt_fast_eval.sh \
  work_dirs/<exp>/iter_XXXX.pth 8

Predicted masks in the model output are split into <think> (reasoning) and <answer> segments and matched against human_labeled_r_objs / human_labeled_a_objs via Hungarian assignment on IoU.

Evaluating closed-source LLMs (Gemini / GPT / OpenRouter-hosted Qwen)

For models without open weights, we provide an OpenRouter + SAM2 evaluator:

📁 projects/vrt_sa2va/eval_openrouter/ — VER mask evaluation via OpenRouter

It (1) prompts the LLM through OpenRouter for bounding boxes, (2) feeds the boxes into SAM2 to produce masks, then (3) scores those masks against the human-labelled GT in this dataset using the same metrics as vrt_eval_single.py.

Metrics

MetricMeaning
answer_miouMain metric. Mean IoU on the answer mask(s)
reasoning_miouMean IoU on the reasoning mask(s)
LQLogical Quality — % of predictions with IoU > 0.5
SQSegmentation Quality — mean IoU among predictions with IoU > 0.5

Question prompt

Match the training prompt for fair comparison:

<image>
{question}

You should first think about the reasoning process in the mind and then
provides the user with the answer. Please respond with segmentation mask in
both the thinking process and the answer.

The reasoning process and answer are enclosed within <think> </think> and
<answer> </answer> tags, respectively, i.e., <think> reasoning process here
</think> <answer> answers here </answer>.

License

Annotations are released under CC BY-NC 4.0. Source images come from SA-1B and are subject to the SA-1B license.

Contributors

HarborYuan

6 commits