Awesome paper for multi-modal llm with grounding ability
21
25 commits
updated Oct 11, 2025
A curated list of Multimodal Large Language Models (or Large Vision Language Model) with grounding ability.
| Format | Desc | Paper |
|---|---|---|
| Decoder on latent | leverage a decoder to ground | PerceptionGPT, NExT-Chat, PSALM, PixelLM, u-LLaVA, GSVA, ChatterBox, GLaMM |
| Output numerical coordinates | direct output numerical tokens | Shikra, VisionLLM, Ferret, Ferret2, CogVLM |
| Output token coordinates | output new tokens added to refer positions | Kosmos-2 |
| Pixel space | output in discrete pixel space encoded by VQGAN | Unified-IO, Unified-IO 2 |
| Proposal retrieval | retrieval from region candidates | LLM-Seg, Kosmos-2, GROUNDHOG |
| Format | Desc | Paper |
|---|---|---|
| Pooling | Leverage Mask Pooling / RoI Pooling / RoI Align to obtain features from the im encoder output | Groma, GPT4RoI, Osprey, PSALM, GROUNDHOG, Ferret, Ferret2, PVIT, ChatterBox, GLaMM |
| Numerical coordinates | Leverage numerical coordinates for referring (bbox / sampled points in mask) | Shikra, PerceptionGPT (w/ encoder), NExT-Chat (w/ encoder), CogVLM |
| Token coordinates | Add new tokens to vocab to present spatial positions | Kosmos-2 |
| Dataset | Source | Data Source | Quantity | Cnstruction Method |
|---|---|---|---|---|
| GRIT | Ferret | COYO-700M, LAION-2B | - | |
| Shikra-RD | Shikra | Flickr30K Entities | 5,922 QA pairs | ChatGPT4 ==> Referential Dialogue (CoT dialogues with grounding & referring) |
| CB-300K | ChatterBox | VG | 717,075 QA pairs | 4 subsets. |
| GranD | GLaMM | SA-1B | 11M images with 7.5M unique concepts and 810M regions. | Automated annotation pipeline with SAM for dense pixel-wise grounding. Used for pretraining. |
| GranD-f | GLaMM | GranD (refined), Flickr30K, RefCOCOg, and PSG | ~214K image-grounded text pairs | Refined subset of GranD for fine-tuning, with 1000 images held out for human-annotated evaluation |
| Model | Recipe |
|---|---|
| Ferret | |
| Ferret2 | |
| ChatterBox | Trainable: LoRA and location decoder |
| GPT4RoI | |
| GLaMM |
| Dataset | Source | Data Source | Quantity | Cnstruction Method |
|---|---|---|---|---|
| Ferret Bench | Ferret | COCO validation set | 120 |




Awesome paper for multi-modal llm with grounding ability
21
25 commits
updated Oct 11, 2025
A curated list of Multimodal Large Language Models (or Large Vision Language Model) with grounding ability.
| Format | Desc | Paper |
|---|---|---|
| Decoder on latent | leverage a decoder to ground | PerceptionGPT, NExT-Chat, PSALM, PixelLM, u-LLaVA, GSVA, ChatterBox, GLaMM |
| Output numerical coordinates | direct output numerical tokens | Shikra, VisionLLM, Ferret, Ferret2, CogVLM |
| Output token coordinates | output new tokens added to refer positions | Kosmos-2 |
| Pixel space | output in discrete pixel space encoded by VQGAN | Unified-IO, Unified-IO 2 |
| Proposal retrieval | retrieval from region candidates | LLM-Seg, Kosmos-2, GROUNDHOG |
| Format | Desc | Paper |
|---|---|---|
| Pooling | Leverage Mask Pooling / RoI Pooling / RoI Align to obtain features from the im encoder output | Groma, GPT4RoI, Osprey, PSALM, GROUNDHOG, Ferret, Ferret2, PVIT, ChatterBox, GLaMM |
| Numerical coordinates | Leverage numerical coordinates for referring (bbox / sampled points in mask) | Shikra, PerceptionGPT (w/ encoder), NExT-Chat (w/ encoder), CogVLM |
| Token coordinates | Add new tokens to vocab to present spatial positions | Kosmos-2 |
| Dataset | Source | Data Source | Quantity | Cnstruction Method |
|---|---|---|---|---|
| GRIT | Ferret | COYO-700M, LAION-2B | - | |
| Shikra-RD | Shikra | Flickr30K Entities | 5,922 QA pairs | ChatGPT4 ==> Referential Dialogue (CoT dialogues with grounding & referring) |
| CB-300K | ChatterBox | VG | 717,075 QA pairs | 4 subsets. |
| GranD | GLaMM | SA-1B | 11M images with 7.5M unique concepts and 810M regions. | Automated annotation pipeline with SAM for dense pixel-wise grounding. Used for pretraining. |
| GranD-f | GLaMM | GranD (refined), Flickr30K, RefCOCOg, and PSG | ~214K image-grounded text pairs | Refined subset of GranD for fine-tuning, with 1000 images held out for human-annotated evaluation |
| Model | Recipe |
|---|---|
| Ferret | |
| Ferret2 | |
| ChatterBox | Trainable: LoRA and location decoder |
| GPT4RoI | |
| GLaMM |
| Dataset | Source | Data Source | Quantity | Cnstruction Method |
|---|---|---|---|---|
| Ferret Bench | Ferret | COCO validation set | 120 |



