This repository only contains the HumanRef Benchmark and the evaluation code.
2
7 commits
1 linked in READMEs
updated May 14, 2025
This repository only contains the HumanRef Benchmark and the evaluation code.
HumanRef is a large-scale human-centric referring expression dataset designed for multi-instance human referring in natural scenes. Unlike traditional referring datasets that focus on one-to-one object referring, HumanRef supports referring to multiple individuals simultaneously through natural language descriptions.
Key features of HumanRef include:
The dataset aims to advance research in human-centric visual understanding and referring expression comprehension in complex, multi-person scenarios.
| Type | Attribute | Position | Interaction | Reasoning | Celebrity | Rejection | Total |
|---|---|---|---|---|---|---|---|
| HumanRef Train | |||||||
| Images | 8,614 | 7,577 | 1,632 | 4,474 | 4,990 | 7,519 | 34,806 |
| Referrings | 52,513 | 22,496 | 2,911 | 6,808 | 4,990 | 13,310 | 103,028 |
| Avg. boxes/ref | 2.9 | 1.9 | 3.1 | 3.0 | 1.0 | 0 | 2.2 |
| HumanRef Benchmark | |||||||
| Images | 838 | 972 | 940 | 982 | 1,000 | 1,000 | 5,732 |
| Referrings | 1,000 | 1,000 | 1,000 | 1,000 | 1,000 | 1,000 | 6,000 |
| Avg. boxes/ref | 2.8 | 2.1 | 2.1 | 2.7 | 1.1 | 0 | 2.2 |
| Dataset | Images | Refs | Vocabs | Avg. Size | Avg. Person/Image | Avg. Words/Ref | Avg. Boxes/Ref |
|---|---|---|---|---|---|---|---|
| RefCOCO | 1,519 | 10,771 | 1,874 | 593x484 | 5.72 | 3.43 | 1 |
| RefCOCO+ | 1,519 | 10,908 | 2,288 | 592x484 | 5.72 | 3.34 | 1 |
| RefCOCOg | 1,521 | 5,253 | 2,479 | 585x480 | 2.73 | 9.07 | 1 |
| HumanRef | 5,732 | 6,000 | 2,714 | 1432x1074 | 8.60 | 6.69 | 2.2 |
Note: For a fair comparison, the statistics for RefCOCO/+/g only include human-referring cases.
HumanRef Benchmark contains 6 domains, each domain may have multiple sub-domains.
| Domain | Subdomain | Num Referrings |
|---|---|---|
| attribute | 1000_attribute_retranslated_with_mask | 1000 |
| position | 500_inner_position_data_with_mask | 500 |
| position | 500_outer_position_data_with_mask | 500 |
| celebrity | 1000_celebrity_data_with_mask | 1000 |
| interaction | 500_inner_interaction_data_with_mask | 500 |
| interaction | 500_outer_interaction_data_with_mask | 500 |
| reasoning | 229_outer_position_two_stage_with_mask | 229 |
| reasoning | 271_positive_then_negative_reasoning_with_mask | 271 |
| reasoning | 500_inner_position_two_stage_with_mask | 500 |
| rejection | 1000_rejection_referring_with_mask | 1000 |
To visualize the dataset, you can run the following command:
python tools/visualize.py \
--anno_path annotations.jsonl \
--image_root_dir images \
--domain_anme attribute \
--sub_domain_anme 1000_attribute_retranslated_with_mask \
--vis_path visualize \
--num_images 50 \
--vis_mask True
We evaluate the referring task using three main metrics: Precision, Recall, and DensityF1 Score.
Precision & Recall: For each referring expression, a predicted bounding box is considered correct if its IoU with any ground truth box exceeds a threshold. Following COCO evaluation protocol, we report average performance across IoU thresholds from 0.5 to 0.95 in steps of 0.05.
Point-based Evaluation: For models that only output points (e.g., Molmo), a prediction is considered correct if the predicted point falls within the mask of the corresponding instance. Note that this is less strict than IoU-based metrics.
Rejection Accuracy: For the rejection subset, we calculate:
Rejection Accuracy = Number of correctly rejected expressions / Total number of expressions
where a correct rejection means the model predicts no boxes for a non-existent reference.
To penalize over-detection (predicting too many boxes), we introduce the DensityF1 Score:
DensityF1 = (1/N) * Σ [2 * (Precision_i * Recall_i)/(Precision_i + Recall_i) * D_i]
where D_i is the density penalty factor:
D_i = min(1.0, GT_Count_i / Predicted_Count_i)
where:
This penalty factor reduces the score when models predict significantly more boxes than the actual number of people in the image, discouraging over-detection strategies.
Before running the evaluation, you need to prepare your model's predictions in the correct format. Each prediction should be a JSON line in a JSONL file with the following structure:
{
"id": "image_id",
"extracted_predictions": [[x1, y1, x2, y2], [x1, y1, x2, y2], ...]
}
Where:
For rejection cases (where no humans should be detected), you should either:
You can run the evaluation script using the following command:
python metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path path/to/your/predictions.jsonl \
--pred_names "Your Model Name" \
--dump_path IDEA-Research/HumanRef/evaluation_results/your_model_results
Parameters:
Evaluating Multiple Models:
To compare multiple models, provide multiple prediction files:
python metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path model1_results.jsonl model2_results.jsonl model3_results.jsonl \
--pred_names "Model 1" "Model 2" "Model 3" \
--dump_path IDEA-Research/HumanRef/evaluation_results/comparison
from metric.recall_precision_densityf1 import recall_precision_densityf1
recall_precision_densityf1(
gt_path="IDEA-Research/HumanRef/annotations.jsonl",
pred_path=["path/to/your/predictions.jsonl"],
dump_path="IDEA-Research/HumanRef/evaluation_results/your_model_results"
)
The evaluation produces several metrics:
The results are broken down by:
The DensityF1 metric is particularly important as it accounts for both precision/recall and the density of humans in the image.
The evaluation generates two tables:
We provide the evaluation results of several models on HumanRef in the evaluation_results folder.
You can also run the evaluation script to compare your model with others.
python metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path \
"IDEA-Research/HumanRef/evaluation_results/eval_deepseekvl2/deepseekvl2_small_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_ferret/ferret7b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_groma/groma7b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_internvl2/internvl2.5_8b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_shikra/shikra7b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_molmo/molmo-7b-d-0924_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_qwen2vl/qwen2.5-7B.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_chatrex/ChatRex-Vicuna7B.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_dinox/dinox_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_rexseek/rexseek_7b.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_full_gt_person/results.jsonl" \
--pred_names \
"DeepSeek-VL2-small" \
"Ferret-7B" \
"Groma-7B" \
"InternVl-2.5-8B" \
"Shikra-7B" \
"Molmo-7B-D-0924" \
"Qwen2.5-VL-7B" \
"ChatRex-7B" \
"DINOX" \
"RexSeek-7B" \
"Baseline" \
--dump_path IDEA-Research/HumanRef/evaluation_results/all_models_comparison
This repository only contains the HumanRef Benchmark and the evaluation code.
2
7 commits
1 linked in READMEs
updated May 14, 2025
This repository only contains the HumanRef Benchmark and the evaluation code.
HumanRef is a large-scale human-centric referring expression dataset designed for multi-instance human referring in natural scenes. Unlike traditional referring datasets that focus on one-to-one object referring, HumanRef supports referring to multiple individuals simultaneously through natural language descriptions.
Key features of HumanRef include:
The dataset aims to advance research in human-centric visual understanding and referring expression comprehension in complex, multi-person scenarios.
| Type | Attribute | Position | Interaction | Reasoning | Celebrity | Rejection | Total |
|---|---|---|---|---|---|---|---|
| HumanRef Train | |||||||
| Images | 8,614 | 7,577 | 1,632 | 4,474 | 4,990 | 7,519 | 34,806 |
| Referrings | 52,513 | 22,496 | 2,911 | 6,808 | 4,990 | 13,310 | 103,028 |
| Avg. boxes/ref | 2.9 | 1.9 | 3.1 | 3.0 | 1.0 | 0 | 2.2 |
| HumanRef Benchmark | |||||||
| Images | 838 | 972 | 940 | 982 | 1,000 | 1,000 | 5,732 |
| Referrings | 1,000 | 1,000 | 1,000 | 1,000 | 1,000 | 1,000 | 6,000 |
| Avg. boxes/ref | 2.8 | 2.1 | 2.1 | 2.7 | 1.1 | 0 | 2.2 |
| Dataset | Images | Refs | Vocabs | Avg. Size | Avg. Person/Image | Avg. Words/Ref | Avg. Boxes/Ref |
|---|---|---|---|---|---|---|---|
| RefCOCO | 1,519 | 10,771 | 1,874 | 593x484 | 5.72 | 3.43 | 1 |
| RefCOCO+ | 1,519 | 10,908 | 2,288 | 592x484 | 5.72 | 3.34 | 1 |
| RefCOCOg | 1,521 | 5,253 | 2,479 | 585x480 | 2.73 | 9.07 | 1 |
| HumanRef | 5,732 | 6,000 | 2,714 | 1432x1074 | 8.60 | 6.69 | 2.2 |
Note: For a fair comparison, the statistics for RefCOCO/+/g only include human-referring cases.
HumanRef Benchmark contains 6 domains, each domain may have multiple sub-domains.
| Domain | Subdomain | Num Referrings |
|---|---|---|
| attribute | 1000_attribute_retranslated_with_mask | 1000 |
| position | 500_inner_position_data_with_mask | 500 |
| position | 500_outer_position_data_with_mask | 500 |
| celebrity | 1000_celebrity_data_with_mask | 1000 |
| interaction | 500_inner_interaction_data_with_mask | 500 |
| interaction | 500_outer_interaction_data_with_mask | 500 |
| reasoning | 229_outer_position_two_stage_with_mask | 229 |
| reasoning | 271_positive_then_negative_reasoning_with_mask | 271 |
| reasoning | 500_inner_position_two_stage_with_mask | 500 |
| rejection | 1000_rejection_referring_with_mask | 1000 |
To visualize the dataset, you can run the following command:
python tools/visualize.py \
--anno_path annotations.jsonl \
--image_root_dir images \
--domain_anme attribute \
--sub_domain_anme 1000_attribute_retranslated_with_mask \
--vis_path visualize \
--num_images 50 \
--vis_mask True
We evaluate the referring task using three main metrics: Precision, Recall, and DensityF1 Score.
Precision & Recall: For each referring expression, a predicted bounding box is considered correct if its IoU with any ground truth box exceeds a threshold. Following COCO evaluation protocol, we report average performance across IoU thresholds from 0.5 to 0.95 in steps of 0.05.
Point-based Evaluation: For models that only output points (e.g., Molmo), a prediction is considered correct if the predicted point falls within the mask of the corresponding instance. Note that this is less strict than IoU-based metrics.
Rejection Accuracy: For the rejection subset, we calculate:
Rejection Accuracy = Number of correctly rejected expressions / Total number of expressions
where a correct rejection means the model predicts no boxes for a non-existent reference.
To penalize over-detection (predicting too many boxes), we introduce the DensityF1 Score:
DensityF1 = (1/N) * Σ [2 * (Precision_i * Recall_i)/(Precision_i + Recall_i) * D_i]
where D_i is the density penalty factor:
D_i = min(1.0, GT_Count_i / Predicted_Count_i)
where:
This penalty factor reduces the score when models predict significantly more boxes than the actual number of people in the image, discouraging over-detection strategies.
Before running the evaluation, you need to prepare your model's predictions in the correct format. Each prediction should be a JSON line in a JSONL file with the following structure:
{
"id": "image_id",
"extracted_predictions": [[x1, y1, x2, y2], [x1, y1, x2, y2], ...]
}
Where:
For rejection cases (where no humans should be detected), you should either:
You can run the evaluation script using the following command:
python metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path path/to/your/predictions.jsonl \
--pred_names "Your Model Name" \
--dump_path IDEA-Research/HumanRef/evaluation_results/your_model_results
Parameters:
Evaluating Multiple Models:
To compare multiple models, provide multiple prediction files:
python metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path model1_results.jsonl model2_results.jsonl model3_results.jsonl \
--pred_names "Model 1" "Model 2" "Model 3" \
--dump_path IDEA-Research/HumanRef/evaluation_results/comparison
from metric.recall_precision_densityf1 import recall_precision_densityf1
recall_precision_densityf1(
gt_path="IDEA-Research/HumanRef/annotations.jsonl",
pred_path=["path/to/your/predictions.jsonl"],
dump_path="IDEA-Research/HumanRef/evaluation_results/your_model_results"
)
The evaluation produces several metrics:
The results are broken down by:
The DensityF1 metric is particularly important as it accounts for both precision/recall and the density of humans in the image.
The evaluation generates two tables:
We provide the evaluation results of several models on HumanRef in the evaluation_results folder.
You can also run the evaluation script to compare your model with others.
python metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path \
"IDEA-Research/HumanRef/evaluation_results/eval_deepseekvl2/deepseekvl2_small_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_ferret/ferret7b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_groma/groma7b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_internvl2/internvl2.5_8b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_shikra/shikra7b_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_molmo/molmo-7b-d-0924_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_qwen2vl/qwen2.5-7B.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_chatrex/ChatRex-Vicuna7B.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_dinox/dinox_results.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_rexseek/rexseek_7b.jsonl" \
"IDEA-Research/HumanRef/evaluation_results/eval_full_gt_person/results.jsonl" \
--pred_names \
"DeepSeek-VL2-small" \
"Ferret-7B" \
"Groma-7B" \
"InternVl-2.5-8B" \
"Shikra-7B" \
"Molmo-7B-D-0924" \
"Qwen2.5-VL-7B" \
"ChatRex-7B" \
"DINOX" \
"RexSeek-7B" \
"Baseline" \
--dump_path IDEA-Research/HumanRef/evaluation_results/all_models_comparison