[ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark
See the code
🔥 [2025/10/15] Rex-Omni: Still using traditional detectors? We've turned object detection into a simple "Next-Token Prediction" task with an MLLM! One model (fully open-sourced), zero-shot SOTA performance, tackling detection, referring, OCR, and GUI grounding all at once. Come see the next generation of perception models
RexSeek is a Multimodal Large Language Model (MLLM) designed to detect people or objects in images based on natural language descriptions. Unlike traditional referring models that focus on single-instance detection, RexSeek excels at multi-instance referring tasks - identifying multiple people or objects that match a given description.
We also introduce HumanRef Benchmark, a comprehensive benchmark for human-centric referring tasks containing:
conda create -n rexseek CUDA_VISIBLE_DEVICES=0 python=3.10
conda activate rexseek
pip install torch==2.1.2 torchvision==0.16.2 --index-url https://download.pytorch.org/whl/cu121
pip install -v -e .
We provide model checkpoints for RexSeek-3B. You can download the pre-trained models from the following links:
Or you can also using the following command to download the pre-trained models:
# Download RexSeek checkpoint from Hugging Face
git lfs install
git clone https://huggingface.co/IDEA-Research/RexSeek-3B IDEA-Research/RexSeek-3B
To verify the installation, run the following command:
CUDA_VISIBLE_DEVICES=0 python tests/test_local_load.py
If the installation is successful, you will get a visualization image in tests/images folder.
TL;DR: RexSeek needs model to propose object boxes first, then use the LLM to detect the objects.
RexSeek consists of three key components:
Inputs:
Outputs:
<ground>referring text</ground><objects><obj1><obj2>...</objects>
In this example, we will use GroundingDINO to generate object proposals, and then use RexSeek to detect the objects.
cd demos/
git clone https://github.com/IDEA-Research/GroundingDINO.git
cd GroundingDINO
pip install -v -e .
pip install numpy==1.26.4
mkdir weights
wget -q https://github.com/IDEA-Research/GroundingDINO/releases/download/v0.1.0-alpha/groundingdino_swint_ogc.pth -P weights
cd ../../../
CUDA_VISIBLE_DEVICES=0 python demos/rexseek_grounding_dino.py \
--image demos/demo_images/demo1.jpg \
--output demos/demo_images/demo1_result.jpg \
--referring "person that is giving a proposal" \
--objects "person" \
--text-threshold 0.25 \
--box-threshold 0.25
In previous example, we need to explicitly specify object categories (like "person") for GroundingDINO to detect. However, we can make this process more automatic by using Spacy to extract nouns from the question as detection targets.
pip install spacy
CUDA_VISIBLE_DEVICES=0 python -m spacy download en_core_web_sm
CUDA_VISIBLE_DEVICES=0 python demos/rexseek_grounding_dino_spacy.py \
--image demos/demo_images/demo1.jpg \
--output demos/demo_images/demo1_result.jpg \
--referring "person that is giving a proposal" \
--text-threshold 0.25 \
--box-threshold 0.25
In this enhanced version:
--objects parameterIn this example, we will use GroundingDINO to generate object proposals, then use Spacy to extract nouns from the question as detection targets, and finally use SAM to segment the objects.
cd demos/
git clone git@github.com:facebookresearch/segment-anything.git
cd segment-anything
pip install -e .
mkdir weights
wget -q https://dl.fbaipublicfiles.com/segment_anything/sam_vit_h_4b8939.pth -P weights
cd ../../../
CUDA_VISIBLE_DEVICES=0 python demos/rexseek_grounding_dino_spacy_sam.py \
--image demos/demo_images/demo1.jpg \
--output demos/demo_images/demo1_result.jpg \
--referring "person that is giving a proposal" \
--text-threshold 0.25 \
--box-threshold 0.25
We provide a gradio demo for RexSeek + GroundingDINO + SAM. You can run the following command to start the gradio demo:
CUDA_VISIBLE_DEVICES=0 python demos/gradio_demo.py \
--rexseek-path "IDEA-Research/RexSeek-3B" \
--gdino-config "demos/GroundingDINO/groundingdino/config/GroundingDINO_SwinT_OGC.py" \
--gdino-weights "demos/GroundingDINO/weights/groundingdino_swint_ogc.pth" \
--sam-weights "demos/segment-anything/weights/sam_vit_h_4b8939.pth"
HumanRef is a large-scale human-centric referring expression dataset designed for multi-instance human referring in natural scenes. Unlike traditional referring datasets that focus on one-to-one object referring, HumanRef supports referring to multiple individuals simultaneously through natural language descriptions.
Key features of HumanRef include:
The dataset aims to advance research in human-centric visual understanding and referring expression comprehension in complex, multi-person scenarios.
You can download the HumanRef Benchmark at https://huggingface.co/datasets/IDEA-Research/HumanRef.
HumanRef Benchmark contains 6 domains, each domain may have multiple sub-domains.
| Domain | Subdomain | Num Referrings |
|---|---|---|
| attribute | 1000_attribute_retranslated_with_mask | 1000 |
| position | 500_inner_position_data_with_mask | 500 |
| position | 500_outer_position_data_with_mask | 500 |
| celebrity | 1000_celebrity_data_with_mask | 1000 |
| interaction | 500_inner_interaction_data_with_mask | 500 |
| interaction | 500_outer_interaction_data_with_mask | 500 |
| reasoning | 229_outer_position_two_stage_with_mask | 229 |
| reasoning | 271_positive_then_negative_reasoning_with_mask | 271 |
| reasoning | 500_inner_position_two_stage_with_mask | 500 |
| rejection | 1000_rejection_referring_with_mask | 1000 |
To visualize the dataset, you can run the following command:
CUDA_VISIBLE_DEVICES=0 python rexseek/tools/visualize_humanref.py \
--anno_path "IDEA-Research/HumanRef/annotations.jsonl" \
--image_root_dir "IDEA-Research/HumanRef/images" \
--domain_name "attribute" \ # attribute, position, interaction, reasoning, celebrity, rejection
--sub_domain_name "1000_attribute_retranslated_with_mask" \ # 1000_attribute_retranslated_with_mask, 500_inner_position_data_with_mask, 500_outer_position_data_with_mask, 1000_celebrity_data_with_mask, 500_inner_interaction_data_with_mask, 500_outer_interaction_data_with_mask, 229_outer_position_two_stage_with_mask, 271_positive_then_negative_reasoning_with_mask, 500_inner_position_two_stage_with_mask, 1000_rejection_referring_with_mask
--vis_path "IDEA-Research/HumanRef/visualize" \
--num_images 50 \
--vis_mask True # True, False
We evaluate the referring task using three main metrics: Precision, Recall, and DensityF1 Score.
Precision & Recall: For each referring expression, a predicted bounding box is considered correct if its IoU with any ground truth box exceeds a threshold. Following COCO evaluation protocol, we report average performance across IoU thresholds from 0.5 to 0.95 in steps of 0.05.
Point-based Evaluation: For models that only output points (e.g., Molmo), a prediction is considered correct if the predicted point falls within the mask of the corresponding instance. Note that this is less strict than IoU-based metrics.
Rejection Accuracy: For the rejection subset, we calculate:
Rejection Accuracy = Number of correctly rejected expressions / Total number of expressions
where a correct rejection means the model predicts no boxes for a non-existent reference.
To penalize over-detection (predicting too many boxes), we introduce the DensityF1 Score:
DensityF1 = (1/N) * Σ [2 * (Precision_i * Recall_i)/(Precision_i + Recall_i) * D_i]
where D_i is the density penalty factor:
D_i = min(1.0, GT_Count_i / Predicted_Count_i)
where:
This penalty factor reduces the score when models predict significantly more boxes than the actual number of people in the image, discouraging over-detection strategies.
Before running the evaluation, you need to prepare your model's predictions in the correct format. Each prediction should be a JSON line in a JSONL file with the following structure:
{
"id": "image_id",
"extracted_predictions": [[x1, y1, x2, y2], [x1, y1, x2, y2], ...]
}
Where:
For rejection cases (where no humans should be detected), you should either:
You can run the evaluation script using the following command:
CUDA_VISIBLE_DEVICES=0 python rexseek/metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path path/to/your/predictions.jsonl \
--pred_names "Your Model Name" \
--dump_path IDEA-Research/HumanRef/evaluation_results/your_model_results
Parameters:
Evaluating Multiple Models:
To compare multiple models, provide multiple prediction files:
CUDA_VISIBLE_DEVICES=0 python rexseek/metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path model1_results.jsonl model2_results.jsonl model3_results.jsonl \
--pred_names "Model 1" "Model 2" "Model 3" \
--dump_path IDEA-Research/HumanRef/evaluation_results/comparison
from rexseek.metric.recall_precision_densityf1 import recall_precision_densityf1
recall_precision_densityf1(
gt_path="IDEA-Research/HumanRef/annotations.jsonl",
pred_path=["path/to/your/predictions.jsonl"],
dump_path="IDEA-Research/HumanRef/evaluation_results/your_model_results"
)
First we need to run the following command to generate the predictions:
CUDA_VISIBLE_DEVICES=0 python rexseek/evaluation/evaluate_rexseek.py \
--model_path IDEA-Research/RexSeek-3B \
--image_folder IDEA-Research/HumanRef/images \
--question_file IDEA-Research/HumanRef/annotations.jsonl \
--answers_file IDEA-Research/HumanRef/evaluation_results/eval_rexseek/RexSeek-3B_results.jsonl \
Then we can run the following command to evaluate the RexSeek model:
python rexseek/metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path IDEA-Research/HumanRef/evaluation_results/eval_rexseek/RexSeek-3B_results.jsonl\
--pred_names "RexSeek-3B" \
--dump_path IDEA-Research/HumanRef/evaluation_results/comparison
45K data with CoT annotations are available at https://huggingface.co/datasets/IDEA-Research/HumanRef-45K.
RexSeek is licensed under the IDEA License 1.0, Copyright (c) IDEA. All Rights Reserved. Note that this project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses including but not limited to the:
@misc{jiang2025referringperson,
title={Referring to Any Person},
author={Qing Jiang and Lin Wu and Zhaoyang Zeng and Tianhe Ren and Yuda Xiong and Yihao Chen and Qin Liu and Lei Zhang},
year={2025},
eprint={2503.08507},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.08507},
}
Python
100.0%
[ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark
See the code
🔥 [2025/10/15] Rex-Omni: Still using traditional detectors? We've turned object detection into a simple "Next-Token Prediction" task with an MLLM! One model (fully open-sourced), zero-shot SOTA performance, tackling detection, referring, OCR, and GUI grounding all at once. Come see the next generation of perception models
RexSeek is a Multimodal Large Language Model (MLLM) designed to detect people or objects in images based on natural language descriptions. Unlike traditional referring models that focus on single-instance detection, RexSeek excels at multi-instance referring tasks - identifying multiple people or objects that match a given description.
We also introduce HumanRef Benchmark, a comprehensive benchmark for human-centric referring tasks containing:
conda create -n rexseek CUDA_VISIBLE_DEVICES=0 python=3.10
conda activate rexseek
pip install torch==2.1.2 torchvision==0.16.2 --index-url https://download.pytorch.org/whl/cu121
pip install -v -e .
We provide model checkpoints for RexSeek-3B. You can download the pre-trained models from the following links:
Or you can also using the following command to download the pre-trained models:
# Download RexSeek checkpoint from Hugging Face
git lfs install
git clone https://huggingface.co/IDEA-Research/RexSeek-3B IDEA-Research/RexSeek-3B
To verify the installation, run the following command:
CUDA_VISIBLE_DEVICES=0 python tests/test_local_load.py
If the installation is successful, you will get a visualization image in tests/images folder.
TL;DR: RexSeek needs model to propose object boxes first, then use the LLM to detect the objects.
RexSeek consists of three key components:
Inputs:
Outputs:
<ground>referring text</ground><objects><obj1><obj2>...</objects>
In this example, we will use GroundingDINO to generate object proposals, and then use RexSeek to detect the objects.
cd demos/
git clone https://github.com/IDEA-Research/GroundingDINO.git
cd GroundingDINO
pip install -v -e .
pip install numpy==1.26.4
mkdir weights
wget -q https://github.com/IDEA-Research/GroundingDINO/releases/download/v0.1.0-alpha/groundingdino_swint_ogc.pth -P weights
cd ../../../
CUDA_VISIBLE_DEVICES=0 python demos/rexseek_grounding_dino.py \
--image demos/demo_images/demo1.jpg \
--output demos/demo_images/demo1_result.jpg \
--referring "person that is giving a proposal" \
--objects "person" \
--text-threshold 0.25 \
--box-threshold 0.25
In previous example, we need to explicitly specify object categories (like "person") for GroundingDINO to detect. However, we can make this process more automatic by using Spacy to extract nouns from the question as detection targets.
pip install spacy
CUDA_VISIBLE_DEVICES=0 python -m spacy download en_core_web_sm
CUDA_VISIBLE_DEVICES=0 python demos/rexseek_grounding_dino_spacy.py \
--image demos/demo_images/demo1.jpg \
--output demos/demo_images/demo1_result.jpg \
--referring "person that is giving a proposal" \
--text-threshold 0.25 \
--box-threshold 0.25
In this enhanced version:
--objects parameterIn this example, we will use GroundingDINO to generate object proposals, then use Spacy to extract nouns from the question as detection targets, and finally use SAM to segment the objects.
cd demos/
git clone git@github.com:facebookresearch/segment-anything.git
cd segment-anything
pip install -e .
mkdir weights
wget -q https://dl.fbaipublicfiles.com/segment_anything/sam_vit_h_4b8939.pth -P weights
cd ../../../
CUDA_VISIBLE_DEVICES=0 python demos/rexseek_grounding_dino_spacy_sam.py \
--image demos/demo_images/demo1.jpg \
--output demos/demo_images/demo1_result.jpg \
--referring "person that is giving a proposal" \
--text-threshold 0.25 \
--box-threshold 0.25
We provide a gradio demo for RexSeek + GroundingDINO + SAM. You can run the following command to start the gradio demo:
CUDA_VISIBLE_DEVICES=0 python demos/gradio_demo.py \
--rexseek-path "IDEA-Research/RexSeek-3B" \
--gdino-config "demos/GroundingDINO/groundingdino/config/GroundingDINO_SwinT_OGC.py" \
--gdino-weights "demos/GroundingDINO/weights/groundingdino_swint_ogc.pth" \
--sam-weights "demos/segment-anything/weights/sam_vit_h_4b8939.pth"
HumanRef is a large-scale human-centric referring expression dataset designed for multi-instance human referring in natural scenes. Unlike traditional referring datasets that focus on one-to-one object referring, HumanRef supports referring to multiple individuals simultaneously through natural language descriptions.
Key features of HumanRef include:
The dataset aims to advance research in human-centric visual understanding and referring expression comprehension in complex, multi-person scenarios.
You can download the HumanRef Benchmark at https://huggingface.co/datasets/IDEA-Research/HumanRef.
HumanRef Benchmark contains 6 domains, each domain may have multiple sub-domains.
| Domain | Subdomain | Num Referrings |
|---|---|---|
| attribute | 1000_attribute_retranslated_with_mask | 1000 |
| position | 500_inner_position_data_with_mask | 500 |
| position | 500_outer_position_data_with_mask | 500 |
| celebrity | 1000_celebrity_data_with_mask | 1000 |
| interaction | 500_inner_interaction_data_with_mask | 500 |
| interaction | 500_outer_interaction_data_with_mask | 500 |
| reasoning | 229_outer_position_two_stage_with_mask | 229 |
| reasoning | 271_positive_then_negative_reasoning_with_mask | 271 |
| reasoning | 500_inner_position_two_stage_with_mask | 500 |
| rejection | 1000_rejection_referring_with_mask | 1000 |
To visualize the dataset, you can run the following command:
CUDA_VISIBLE_DEVICES=0 python rexseek/tools/visualize_humanref.py \
--anno_path "IDEA-Research/HumanRef/annotations.jsonl" \
--image_root_dir "IDEA-Research/HumanRef/images" \
--domain_name "attribute" \ # attribute, position, interaction, reasoning, celebrity, rejection
--sub_domain_name "1000_attribute_retranslated_with_mask" \ # 1000_attribute_retranslated_with_mask, 500_inner_position_data_with_mask, 500_outer_position_data_with_mask, 1000_celebrity_data_with_mask, 500_inner_interaction_data_with_mask, 500_outer_interaction_data_with_mask, 229_outer_position_two_stage_with_mask, 271_positive_then_negative_reasoning_with_mask, 500_inner_position_two_stage_with_mask, 1000_rejection_referring_with_mask
--vis_path "IDEA-Research/HumanRef/visualize" \
--num_images 50 \
--vis_mask True # True, False
We evaluate the referring task using three main metrics: Precision, Recall, and DensityF1 Score.
Precision & Recall: For each referring expression, a predicted bounding box is considered correct if its IoU with any ground truth box exceeds a threshold. Following COCO evaluation protocol, we report average performance across IoU thresholds from 0.5 to 0.95 in steps of 0.05.
Point-based Evaluation: For models that only output points (e.g., Molmo), a prediction is considered correct if the predicted point falls within the mask of the corresponding instance. Note that this is less strict than IoU-based metrics.
Rejection Accuracy: For the rejection subset, we calculate:
Rejection Accuracy = Number of correctly rejected expressions / Total number of expressions
where a correct rejection means the model predicts no boxes for a non-existent reference.
To penalize over-detection (predicting too many boxes), we introduce the DensityF1 Score:
DensityF1 = (1/N) * Σ [2 * (Precision_i * Recall_i)/(Precision_i + Recall_i) * D_i]
where D_i is the density penalty factor:
D_i = min(1.0, GT_Count_i / Predicted_Count_i)
where:
This penalty factor reduces the score when models predict significantly more boxes than the actual number of people in the image, discouraging over-detection strategies.
Before running the evaluation, you need to prepare your model's predictions in the correct format. Each prediction should be a JSON line in a JSONL file with the following structure:
{
"id": "image_id",
"extracted_predictions": [[x1, y1, x2, y2], [x1, y1, x2, y2], ...]
}
Where:
For rejection cases (where no humans should be detected), you should either:
You can run the evaluation script using the following command:
CUDA_VISIBLE_DEVICES=0 python rexseek/metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path path/to/your/predictions.jsonl \
--pred_names "Your Model Name" \
--dump_path IDEA-Research/HumanRef/evaluation_results/your_model_results
Parameters:
Evaluating Multiple Models:
To compare multiple models, provide multiple prediction files:
CUDA_VISIBLE_DEVICES=0 python rexseek/metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path model1_results.jsonl model2_results.jsonl model3_results.jsonl \
--pred_names "Model 1" "Model 2" "Model 3" \
--dump_path IDEA-Research/HumanRef/evaluation_results/comparison
from rexseek.metric.recall_precision_densityf1 import recall_precision_densityf1
recall_precision_densityf1(
gt_path="IDEA-Research/HumanRef/annotations.jsonl",
pred_path=["path/to/your/predictions.jsonl"],
dump_path="IDEA-Research/HumanRef/evaluation_results/your_model_results"
)
First we need to run the following command to generate the predictions:
CUDA_VISIBLE_DEVICES=0 python rexseek/evaluation/evaluate_rexseek.py \
--model_path IDEA-Research/RexSeek-3B \
--image_folder IDEA-Research/HumanRef/images \
--question_file IDEA-Research/HumanRef/annotations.jsonl \
--answers_file IDEA-Research/HumanRef/evaluation_results/eval_rexseek/RexSeek-3B_results.jsonl \
Then we can run the following command to evaluate the RexSeek model:
python rexseek/metric/recall_precision_densityf1.py \
--gt_path IDEA-Research/HumanRef/annotations.jsonl \
--pred_path IDEA-Research/HumanRef/evaluation_results/eval_rexseek/RexSeek-3B_results.jsonl\
--pred_names "RexSeek-3B" \
--dump_path IDEA-Research/HumanRef/evaluation_results/comparison
45K data with CoT annotations are available at https://huggingface.co/datasets/IDEA-Research/HumanRef-45K.
RexSeek is licensed under the IDEA License 1.0, Copyright (c) IDEA. All Rights Reserved. Note that this project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses including but not limited to the:
@misc{jiang2025referringperson,
title={Referring to Any Person},
author={Qing Jiang and Lin Wu and Zhaoyang Zeng and Tianhe Ren and Yuda Xiong and Yihao Chen and Qin Liu and Lei Zhang},
year={2025},
eprint={2503.08507},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.08507},
}
Python
100.0%