IDEA-Research/HumanRef

Dataset

This repository only contains the HumanRef Benchmark and the evaluation code.

2

7 commits

1 linked in READMEs

updated May 14, 2025

See the code

README

This repository only contains the HumanRef Benchmark and the evaluation code.

1. Introduction

HumanRef is a large-scale human-centric referring expression dataset designed for multi-instance human referring in natural scenes. Unlike traditional referring datasets that focus on one-to-one object referring, HumanRef supports referring to multiple individuals simultaneously through natural language descriptions.

Key features of HumanRef include:

  • Multi-Instance Referring: A single referring expression can correspond to multiple individuals, better reflecting real-world scenarios
  • Diverse Referring Types: Covers 6 major types of referring expressions:
    • Attribute-based (e.g., gender, age, clothing)
    • Position-based (relative positions between humans or with environment)
    • Interaction-based (human-human or human-environment interactions)
    • Reasoning-based (complex logical combinations)
    • Celebrity Recognition
    • Rejection Cases (non-existent references)
  • High-Quality Data:
    • 34,806 high-resolution images (>1000×1000 pixels)
    • 103,028 referring expressions in training set
    • 6,000 carefully curated expressions in benchmark set
    • Average 8.6 persons per image
    • Average 2.2 target boxes per referring expression

The dataset aims to advance research in human-centric visual understanding and referring expression comprehension in complex, multi-person scenarios.

2. Statistics

HumanRef Dataset Statistics

TypeAttributePositionInteractionReasoningCelebrityRejectionTotal
HumanRef Train
Images8,6147,5771,6324,4744,9907,51934,806
Referrings52,51322,4962,9116,8084,99013,310103,028
Avg. boxes/ref2.91.93.13.01.002.2
HumanRef Benchmark
Images8389729409821,0001,0005,732
Referrings1,0001,0001,0001,0001,0001,0006,000
Avg. boxes/ref2.82.12.12.71.102.2

Comparison with Existing Datasets

DatasetImagesRefsVocabsAvg. SizeAvg. Person/ImageAvg. Words/RefAvg. Boxes/Ref
RefCOCO1,51910,7711,874593x4845.723.431
RefCOCO+1,51910,9082,288592x4845.723.341
RefCOCOg1,5215,2532,479585x4802.739.071
HumanRef5,7326,0002,7141432x10748.606.692.2

Note: For a fair comparison, the statistics for RefCOCO/+/g only include human-referring cases.

Distribution Visualization

3. Usage

3.1 Visualization

HumanRef Benchmark contains 6 domains, each domain may have multiple sub-domains.

DomainSubdomainNum Referrings
attribute1000_attribute_retranslated_with_mask1000
position500_inner_position_data_with_mask500
position500_outer_position_data_with_mask500
celebrity1000_celebrity_data_with_mask1000
interaction500_inner_interaction_data_with_mask500
interaction500_outer_interaction_data_with_mask500
reasoning229_outer_position_two_stage_with_mask229
reasoning271_positive_then_negative_reasoning_with_mask271
reasoning500_inner_position_two_stage_with_mask500
rejection1000_rejection_referring_with_mask1000

To visualize the dataset, you can run the following command:

python tools/visualize.py \
    --anno_path annotations.jsonl \
    --image_root_dir images \
    --domain_anme attribute \
    --sub_domain_anme 1000_attribute_retranslated_with_mask \
    --vis_path visualize \
    --num_images 50 \
    --vis_mask True 

3.2 Evaluation

3.2.1 Metrics

We evaluate the referring task using three main metrics: Precision, Recall, and DensityF1 Score.

Basic Metrics

  • Precision & Recall: For each referring expression, a predicted bounding box is considered correct if its IoU with any ground truth box exceeds a threshold. Following COCO evaluation protocol, we report average performance across IoU thresholds from 0.5 to 0.95 in steps of 0.05.

  • Point-based Evaluation: For models that only output points (e.g., Molmo), a prediction is considered correct if the predicted point falls within the mask of the corresponding instance. Note that this is less strict than IoU-based metrics.

  • Rejection Accuracy: For the rejection subset, we calculate:

    Rejection Accuracy = Number of correctly rejected expressions / Total number of expressions
    

    where a correct rejection means the model predicts no boxes for a non-existent reference.

DensityF1 Score

To penalize over-detection (predicting too many boxes), we introduce the DensityF1 Score:

DensityF1 = (1/N) * Σ [2 * (Precision_i * Recall_i)/(Precision_i + Recall_i) * D_i]

where D_i is the density penalty factor:

D_i = min(1.0, GT_Count_i / Predicted_Count_i)

where:

  • N is the number of referring expressions
  • GT_Count_i is the total number of persons in image i
  • Predicted_Count_i is the number of predicted boxes for referring expression i

This penalty factor reduces the score when models predict significantly more boxes than the actual number of people in the image, discouraging over-detection strategies.

3.2.2 Evaluation Script

Prediction Format

Before running the evaluation, you need to prepare your model's predictions in the correct format. Each prediction should be a JSON line in a JSONL file with the following structure:

{
  "id": "image_id",
  "extracted_predictions": [[x1, y1, x2, y2], [x1, y1, x2, y2], ...]
}

Where:

  • id: The image identifier matching the ground truth data
  • extracted_predictions: A list of bounding boxes in [x1, y1, x2, y2] format or points in [x, y] format

For rejection cases (where no humans should be detected), you should either:

  • Include an empty list: "extracted_predictions": []
  • Include a list with an empty box: "extracted_predictions": [[]]

Running the Evaluation

You can run the evaluation script using the following command:

python metric/recall_precision_densityf1.py \
  --gt_path IDEA-Research/HumanRef/annotations.jsonl \
  --pred_path path/to/your/predictions.jsonl \
  --pred_names "Your Model Name" \
  --dump_path IDEA-Research/HumanRef/evaluation_results/your_model_results

Parameters:

  • --gt_path: Path to the ground truth annotations file
  • --pred_path: Path to your prediction file(s). You can provide multiple paths to compare different models
  • --pred_names: Names for your models (for display in the results)
  • --dump_path: Directory to save the evaluation results in markdown and JSON formats

Evaluating Multiple Models:

To compare multiple models, provide multiple prediction files:

python metric/recall_precision_densityf1.py \
  --gt_path IDEA-Research/HumanRef/annotations.jsonl \
  --pred_path model1_results.jsonl model2_results.jsonl model3_results.jsonl \
  --pred_names "Model 1" "Model 2" "Model 3" \
  --dump_path IDEA-Research/HumanRef/evaluation_results/comparison

Programmatic Usage

from metric.recall_precision_densityf1 import recall_precision_densityf1

recall_precision_densityf1(
    gt_path="IDEA-Research/HumanRef/annotations.jsonl",
    pred_path=["path/to/your/predictions.jsonl"],
    dump_path="IDEA-Research/HumanRef/evaluation_results/your_model_results"
)

Metrics Explained

The evaluation produces several metrics:

  1. For point predictions:
    • Recall@Point
    • Precision@Point
    • DensityF1@Point
  2. For box predictions:
    • Recall@0.5 (IoU threshold of 0.5)
    • Recall@0.5:0.95 (mean recall across IoU thresholds from 0.5 to 0.95)
    • Precision@0.5
    • Precision@0.5:0.95
    • DensityF1@0.5
    • DensityF1@0.5:0.95
  3. Rejection Score: Accuracy in correctly identifying images with no humans

The results are broken down by:

  • Domain and subdomain
  • Box count ranges (1, 2-5, 6-10, >10)

The DensityF1 metric is particularly important as it accounts for both precision/recall and the density of humans in the image.

Output

The evaluation generates two tables:

  • Comparative Domain and Subdomain Metrics
  • Comparative Box Count Metrics These are displayed in the console and saved as markdown and JSON files if a dump path is provided.

3.2.3 Comparison with Other Models

We provide the evaluation results of several models on HumanRef in the evaluation_results folder.

You can also run the evaluation script to compare your model with others.

python metric/recall_precision_densityf1.py \
  --gt_path IDEA-Research/HumanRef/annotations.jsonl \
  --pred_path \
    "IDEA-Research/HumanRef/evaluation_results/eval_deepseekvl2/deepseekvl2_small_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_ferret/ferret7b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_groma/groma7b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_internvl2/internvl2.5_8b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_shikra/shikra7b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_molmo/molmo-7b-d-0924_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_qwen2vl/qwen2.5-7B.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_chatrex/ChatRex-Vicuna7B.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_dinox/dinox_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_rexseek/rexseek_7b.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_full_gt_person/results.jsonl" \
  --pred_names \
    "DeepSeek-VL2-small" \
    "Ferret-7B" \
    "Groma-7B" \
    "InternVl-2.5-8B" \
    "Shikra-7B" \
    "Molmo-7B-D-0924" \
    "Qwen2.5-VL-7B" \
    "ChatRex-7B" \
    "DINOX" \
    "RexSeek-7B" \
    "Baseline" \
  --dump_path IDEA-Research/HumanRef/evaluation_results/all_models_comparison

IDEA-Research/HumanRef

Dataset

This repository only contains the HumanRef Benchmark and the evaluation code.

2

7 commits

1 linked in READMEs

updated May 14, 2025

See the code

README

This repository only contains the HumanRef Benchmark and the evaluation code.

1. Introduction

HumanRef is a large-scale human-centric referring expression dataset designed for multi-instance human referring in natural scenes. Unlike traditional referring datasets that focus on one-to-one object referring, HumanRef supports referring to multiple individuals simultaneously through natural language descriptions.

Key features of HumanRef include:

  • Multi-Instance Referring: A single referring expression can correspond to multiple individuals, better reflecting real-world scenarios
  • Diverse Referring Types: Covers 6 major types of referring expressions:
    • Attribute-based (e.g., gender, age, clothing)
    • Position-based (relative positions between humans or with environment)
    • Interaction-based (human-human or human-environment interactions)
    • Reasoning-based (complex logical combinations)
    • Celebrity Recognition
    • Rejection Cases (non-existent references)
  • High-Quality Data:
    • 34,806 high-resolution images (>1000×1000 pixels)
    • 103,028 referring expressions in training set
    • 6,000 carefully curated expressions in benchmark set
    • Average 8.6 persons per image
    • Average 2.2 target boxes per referring expression

The dataset aims to advance research in human-centric visual understanding and referring expression comprehension in complex, multi-person scenarios.

2. Statistics

HumanRef Dataset Statistics

TypeAttributePositionInteractionReasoningCelebrityRejectionTotal
HumanRef Train
Images8,6147,5771,6324,4744,9907,51934,806
Referrings52,51322,4962,9116,8084,99013,310103,028
Avg. boxes/ref2.91.93.13.01.002.2
HumanRef Benchmark
Images8389729409821,0001,0005,732
Referrings1,0001,0001,0001,0001,0001,0006,000
Avg. boxes/ref2.82.12.12.71.102.2

Comparison with Existing Datasets

DatasetImagesRefsVocabsAvg. SizeAvg. Person/ImageAvg. Words/RefAvg. Boxes/Ref
RefCOCO1,51910,7711,874593x4845.723.431
RefCOCO+1,51910,9082,288592x4845.723.341
RefCOCOg1,5215,2532,479585x4802.739.071
HumanRef5,7326,0002,7141432x10748.606.692.2

Note: For a fair comparison, the statistics for RefCOCO/+/g only include human-referring cases.

Distribution Visualization

3. Usage

3.1 Visualization

HumanRef Benchmark contains 6 domains, each domain may have multiple sub-domains.

DomainSubdomainNum Referrings
attribute1000_attribute_retranslated_with_mask1000
position500_inner_position_data_with_mask500
position500_outer_position_data_with_mask500
celebrity1000_celebrity_data_with_mask1000
interaction500_inner_interaction_data_with_mask500
interaction500_outer_interaction_data_with_mask500
reasoning229_outer_position_two_stage_with_mask229
reasoning271_positive_then_negative_reasoning_with_mask271
reasoning500_inner_position_two_stage_with_mask500
rejection1000_rejection_referring_with_mask1000

To visualize the dataset, you can run the following command:

python tools/visualize.py \
    --anno_path annotations.jsonl \
    --image_root_dir images \
    --domain_anme attribute \
    --sub_domain_anme 1000_attribute_retranslated_with_mask \
    --vis_path visualize \
    --num_images 50 \
    --vis_mask True 

3.2 Evaluation

3.2.1 Metrics

We evaluate the referring task using three main metrics: Precision, Recall, and DensityF1 Score.

Basic Metrics

  • Precision & Recall: For each referring expression, a predicted bounding box is considered correct if its IoU with any ground truth box exceeds a threshold. Following COCO evaluation protocol, we report average performance across IoU thresholds from 0.5 to 0.95 in steps of 0.05.

  • Point-based Evaluation: For models that only output points (e.g., Molmo), a prediction is considered correct if the predicted point falls within the mask of the corresponding instance. Note that this is less strict than IoU-based metrics.

  • Rejection Accuracy: For the rejection subset, we calculate:

    Rejection Accuracy = Number of correctly rejected expressions / Total number of expressions
    

    where a correct rejection means the model predicts no boxes for a non-existent reference.

DensityF1 Score

To penalize over-detection (predicting too many boxes), we introduce the DensityF1 Score:

DensityF1 = (1/N) * Σ [2 * (Precision_i * Recall_i)/(Precision_i + Recall_i) * D_i]

where D_i is the density penalty factor:

D_i = min(1.0, GT_Count_i / Predicted_Count_i)

where:

  • N is the number of referring expressions
  • GT_Count_i is the total number of persons in image i
  • Predicted_Count_i is the number of predicted boxes for referring expression i

This penalty factor reduces the score when models predict significantly more boxes than the actual number of people in the image, discouraging over-detection strategies.

3.2.2 Evaluation Script

Prediction Format

Before running the evaluation, you need to prepare your model's predictions in the correct format. Each prediction should be a JSON line in a JSONL file with the following structure:

{
  "id": "image_id",
  "extracted_predictions": [[x1, y1, x2, y2], [x1, y1, x2, y2], ...]
}

Where:

  • id: The image identifier matching the ground truth data
  • extracted_predictions: A list of bounding boxes in [x1, y1, x2, y2] format or points in [x, y] format

For rejection cases (where no humans should be detected), you should either:

  • Include an empty list: "extracted_predictions": []
  • Include a list with an empty box: "extracted_predictions": [[]]

Running the Evaluation

You can run the evaluation script using the following command:

python metric/recall_precision_densityf1.py \
  --gt_path IDEA-Research/HumanRef/annotations.jsonl \
  --pred_path path/to/your/predictions.jsonl \
  --pred_names "Your Model Name" \
  --dump_path IDEA-Research/HumanRef/evaluation_results/your_model_results

Parameters:

  • --gt_path: Path to the ground truth annotations file
  • --pred_path: Path to your prediction file(s). You can provide multiple paths to compare different models
  • --pred_names: Names for your models (for display in the results)
  • --dump_path: Directory to save the evaluation results in markdown and JSON formats

Evaluating Multiple Models:

To compare multiple models, provide multiple prediction files:

python metric/recall_precision_densityf1.py \
  --gt_path IDEA-Research/HumanRef/annotations.jsonl \
  --pred_path model1_results.jsonl model2_results.jsonl model3_results.jsonl \
  --pred_names "Model 1" "Model 2" "Model 3" \
  --dump_path IDEA-Research/HumanRef/evaluation_results/comparison

Programmatic Usage

from metric.recall_precision_densityf1 import recall_precision_densityf1

recall_precision_densityf1(
    gt_path="IDEA-Research/HumanRef/annotations.jsonl",
    pred_path=["path/to/your/predictions.jsonl"],
    dump_path="IDEA-Research/HumanRef/evaluation_results/your_model_results"
)

Metrics Explained

The evaluation produces several metrics:

  1. For point predictions:
    • Recall@Point
    • Precision@Point
    • DensityF1@Point
  2. For box predictions:
    • Recall@0.5 (IoU threshold of 0.5)
    • Recall@0.5:0.95 (mean recall across IoU thresholds from 0.5 to 0.95)
    • Precision@0.5
    • Precision@0.5:0.95
    • DensityF1@0.5
    • DensityF1@0.5:0.95
  3. Rejection Score: Accuracy in correctly identifying images with no humans

The results are broken down by:

  • Domain and subdomain
  • Box count ranges (1, 2-5, 6-10, >10)

The DensityF1 metric is particularly important as it accounts for both precision/recall and the density of humans in the image.

Output

The evaluation generates two tables:

  • Comparative Domain and Subdomain Metrics
  • Comparative Box Count Metrics These are displayed in the console and saved as markdown and JSON files if a dump path is provided.

3.2.3 Comparison with Other Models

We provide the evaluation results of several models on HumanRef in the evaluation_results folder.

You can also run the evaluation script to compare your model with others.

python metric/recall_precision_densityf1.py \
  --gt_path IDEA-Research/HumanRef/annotations.jsonl \
  --pred_path \
    "IDEA-Research/HumanRef/evaluation_results/eval_deepseekvl2/deepseekvl2_small_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_ferret/ferret7b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_groma/groma7b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_internvl2/internvl2.5_8b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_shikra/shikra7b_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_molmo/molmo-7b-d-0924_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_qwen2vl/qwen2.5-7B.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_chatrex/ChatRex-Vicuna7B.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_dinox/dinox_results.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_rexseek/rexseek_7b.jsonl" \
    "IDEA-Research/HumanRef/evaluation_results/eval_full_gt_person/results.jsonl" \
  --pred_names \
    "DeepSeek-VL2-small" \
    "Ferret-7B" \
    "Groma-7B" \
    "InternVl-2.5-8B" \
    "Shikra-7B" \
    "Molmo-7B-D-0924" \
    "Qwen2.5-VL-7B" \
    "ChatRex-7B" \
    "DINOX" \
    "RexSeek-7B" \
    "Baseline" \
  --dump_path IDEA-Research/HumanRef/evaluation_results/all_models_comparison