ParaDLC-Bench (Parallel Detailed Localized Captioning Benchmark) is a benchmark for multi-region localized captioning that jointly evaluates caption quality and inference efficiency. It extends DLC-Bench from single-region evaluation to concurrent multi-region evaluation, explicitly stressing a model's ability to describe many regions at once without cross-region interference.
π Paper Β |Β π» Code Β |Β π€ PerceptionDLM
| Source | Images | Masks |
|---|---|---|
| Objects365 V2 (val) | 54 | 178 |
| DaTaSeg Objects365 Instance Seg. | 46 | 121 |
| Total | 100 | 299 |
The evaluation is a two-step, reference-free process:
temperature=0) scores each caption against predefined questions:
Scores are averaged per mask to ensure equal weighting across regions. The benchmark is robust to the choice of judge (verified across GPT-5.2, Gemini-3.1-Pro, and Qwen3.5-27B).
annotations/
βββ annotations.json # images + region masks
βββ qa.json # positive / negative questions per mask
βββ class_names.json # target category names
| Method | Type | Avg (%) | Time (s) |
|---|---|---|---|
| GAR-8B | AR | 69.5 | 479 |
| LLaDA-V-8B | Diffusion | 35.2 | 3241 |
| PerceptionDLM-8B | Diffusion (parallel) | 62.4 | 276 |
See the Evaluation Guide for the full inference + judging pipeline.
python evaluation/ParaDLC-Bench/infer_mask_captions_paradlc.py \
--model-path MSALab/PerceptionDLM \
--image-root annotations/images \
--anno-json annotations/annotations.json \
--qa-json annotations/qa.json \
--gen-length 32 --steps 32 --temperature 0.0 --top-p 1.0
@article{sun2026perceptiondlm,
title = {PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models},
author = {Sun, Yueyi and Wang, Yuhao and Li, Jason and Tian, Ye and Zhang, Tao and Mai, Jacky and Wang, Yihan and Wang, Haochen and Bai, Jinbin and Yang, Ling and Tong, Yunhai},
journal = {arXiv preprint arXiv:2606.19534},
year = {2026}
}
Released under the Apache License 2.0. Source images originate from Objects365 and DaTaSeg; please also respect their original licenses and terms of use.
ParaDLC-Bench (Parallel Detailed Localized Captioning Benchmark) is a benchmark for multi-region localized captioning that jointly evaluates caption quality and inference efficiency. It extends DLC-Bench from single-region evaluation to concurrent multi-region evaluation, explicitly stressing a model's ability to describe many regions at once without cross-region interference.
π Paper Β |Β π» Code Β |Β π€ PerceptionDLM
| Source | Images | Masks |
|---|---|---|
| Objects365 V2 (val) | 54 | 178 |
| DaTaSeg Objects365 Instance Seg. | 46 | 121 |
| Total | 100 | 299 |
The evaluation is a two-step, reference-free process:
temperature=0) scores each caption against predefined questions:
Scores are averaged per mask to ensure equal weighting across regions. The benchmark is robust to the choice of judge (verified across GPT-5.2, Gemini-3.1-Pro, and Qwen3.5-27B).
annotations/
βββ annotations.json # images + region masks
βββ qa.json # positive / negative questions per mask
βββ class_names.json # target category names
| Method | Type | Avg (%) | Time (s) |
|---|---|---|---|
| GAR-8B | AR | 69.5 | 479 |
| LLaDA-V-8B | Diffusion | 35.2 | 3241 |
| PerceptionDLM-8B | Diffusion (parallel) | 62.4 | 276 |
See the Evaluation Guide for the full inference + judging pipeline.
python evaluation/ParaDLC-Bench/infer_mask_captions_paradlc.py \
--model-path MSALab/PerceptionDLM \
--image-root annotations/images \
--anno-json annotations/annotations.json \
--qa-json annotations/qa.json \
--gen-length 32 --steps 32 --temperature 0.0 --top-p 1.0
@article{sun2026perceptiondlm,
title = {PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models},
author = {Sun, Yueyi and Wang, Yuhao and Li, Jason and Tian, Ye and Zhang, Tao and Mai, Jacky and Wang, Yihan and Wang, Haochen and Bai, Jinbin and Yang, Ling and Tong, Yunhai},
journal = {arXiv preprint arXiv:2606.19534},
year = {2026}
}
Released under the Apache License 2.0. Source images originate from Objects365 and DaTaSeg; please also respect their original licenses and terms of use.