williamium3000/awesome-mllm-grounding

Awesome paper for multi-modal llm with grounding ability

21

25 commits

updated Oct 11, 2025

See the code

README

Awesome-Multimodal-Large-Language-Models-With-Grounding

A curated list of Multimodal Large Language Models (or Large Vision Language Model) with grounding ability.

Table of Contents

🔥 Large Vision-Language Model

Grounding

FormatDescPaper
Decoder on latentleverage a decoder to groundPerceptionGPT, NExT-Chat, PSALM, PixelLM, u-LLaVA, GSVA, ChatterBox, GLaMM
Output numerical coordinatesdirect output numerical tokensShikra, VisionLLM, Ferret, Ferret2, CogVLM
Output token coordinatesoutput new tokens added to refer positionsKosmos-2
Pixel spaceoutput in discrete pixel space encoded by VQGANUnified-IO, Unified-IO 2
Proposal retrievalretrieval from region candidatesLLM-Seg, Kosmos-2, GROUNDHOG

Referring

FormatDescPaper
PoolingLeverage Mask Pooling / RoI Pooling / RoI Align to obtain features from the im encoder outputGroma, GPT4RoI, Osprey, PSALM, GROUNDHOG, Ferret, Ferret2, PVIT, ChatterBox, GLaMM
Numerical coordinatesLeverage numerical coordinates for referring (bbox / sampled points in mask)Shikra, PerceptionGPT (w/ encoder), NExT-Chat (w/ encoder), CogVLM
Token coordinatesAdd new tokens to vocab to present spatial positionsKosmos-2
  • w/ encoder: refers to using a encoder to encode the input coordinates.

Training Dataset

DatasetSourceData SourceQuantityCnstruction Method
GRITFerretCOYO-700M, LAION-2B-
  • Templates are used to convert data.
  • SAM is used to generate masks for free-form referring.
  • ChatGPT4 is used to generate dialogues with bbox.
  • Use GLIPv2 to ground groundable nouns in LLaVA-158k.
  • Negative mining: generate negative yes/or question
  • Shikra-RDShikraFlickr30K Entities5,922 QA pairsChatGPT4 ==> Referential Dialogue (CoT dialogues with grounding & referring)
    CB-300KChatterBoxVG717,075 QA pairs4 subsets.
  • CB-MRG: Use ChatGPT to write dialogues with bbox
  • CB-LC, extend strict relation (from scene graph) to multi-turn QA with ChatGPT
  • CB-REF REG task
  • CB-GND: grounding task
  • GranDGLaMMSA-1B11M images with 7.5M unique concepts and 810M regions.Automated annotation pipeline with SAM for dense pixel-wise grounding. Used for pretraining.
    GranD-fGLaMMGranD (refined), Flickr30K, RefCOCOg, and PSG~214K image-grounded text pairsRefined subset of GranD for fine-tuning, with 1000 images held out for human-annotated evaluation

    Training Recipe

    ModelRecipe
    Ferret
  • Use LLaVA pretrained
  • SFT on GRIT
  • Ferret2
  • image-caption alignment on 1.4M image-text pairs
  • high-resolution dense alignment with template referring & grounding
  • instruction tuning with GRIT, VQA and OCR (VQA and OCR are augmented with GLIPv2 bbox)
  • ChatterBoxTrainable: LoRA and location decoder
  • warm up training with visual grounding only dataset.
  • instruction tuning with CB-300K
  • GPT4RoI
  • Use LLaVA pretrained
  • pretrain region feature extractor with text-region datasets (COCO, RefCOCO, RefCOCO+)
  • train connector, region feature extractor and LLM to follow instructions
  • GLaMM
  • Use GPT4RoI pretrained
  • pretrain on 11M GranD with LoRA
  • finetune on GranD-f, LLaVA-Instruct150K and LLaVA-Instruct-80K
  • Evaluation Dataset

    DatasetSourceData SourceQuantityCnstruction Method
    Ferret BenchFerretCOCO validation set120
  • Referring Description: models are asked to describe a referred region based on its interaction with surrounding objects.
  • Referring Reasoning: models need to reason on top of one or more referred regions correctly.
  • Grounding in Conversation: models are required to reason correctly and accurately ground/localize the objects/regions necessary for the reasoning.
  • Paper List

    GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

    Paper | Github

    1. propose referring for mllm by replacing placeholder <region_i> by feature obtained by mask pooling
    Osprey: Pixel Understanding with Visual Instruction Tuning

    Paper | Github

    1. similar to GPT4RoI, Osprey also use mask representation to refer to entities in images.
    2. It uses mask pooling to extract semantic features from image encoder and combines with a location extractor to process the mask and output spatial token.
    VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

    Paper | Github

    1. unified interface for vision and vl tasks: points for detection, sample points for instance seg ==> instruction format for training
    2. extra tokens & output-format-as-query to decode (faster)
    Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

    Paper | Github | Project

    1. creates a unified IO for all sorts of vision and vl task (into discrete tokens)
    2. using t5-like encoder-decoder arch
    Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

    Paper | Github | Project

    1. following Unified-IO v1, creates a unified IO for all sorts of modalities including image, masks, bboxes, audios (into discrete tokens)
      1. dense masks are all binary, unlike v1 which specifies the color in text instruction (model struggles to follow)
    2. propose 2D Rotary Embedding, QK Normalization and Scaled Cosine Attention to stabilize training and scaling
    3. Mixture of Denoisers taining objectives
    4. instruction tuning of 220 tasks drawn from over 120 external datasets
    PixelLM: Pixel Reasoning with Large Multimodal Model

    Paper | Github | Project

    1. learnable seg tokens + light-weight decoder
    2. a bunch of tricks:
      1. N x L seg tokens for L level multi-scale vision features. N tokens within each group for better modeling
      2. reweighted loss on regions with overlapping predictions
    PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

    Paper | Github

    1. new paradigm: first generate mask proposal, then genereate mask and classification (following mask2former)
    2. instruction prompt + conditional prompt + candidate masks token
      1. three types of conditional prompt: classes, sentence (ref seg) and visual cues (point, scribbles, boxes, etc)
      2. conditional prompt => condition embed, candidate masks token => mask embed.
      3. condition embed +mask embed + image feature => mask2former decoder => bipartite matching loss + query-based decoding 图 0
    LLM-Seg: Bridging Image Segmentation and Large Language Model Reasoning

    Paper | Github

    1. Use SAM to generate mask candidates, then fomulate the problem as mask selection (mask classification)
    2. promote LLM-Seg40K dataset, by using LLaVA to generate caption, then GPT4 to generate question-answer pair.
    GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

    Paper | Project

    1. disantengle grounding with referring
    2. grounding as mask selection and train a mask2former+ to generate mask candidates
    3. referring by mask pooling on feature
    4. promote 2.5M M3G2 dataset
    DetGPT: Detect What You Need via Reasoning

    Paper | Github | Project

    1. Follow LLaVA to tune VLM and for vqa
    2. Use grouding DINO to ground response generated by VLM to detect the relevantg entities.
    Ferret: Refer and Ground Anything Anywhere at Any Granularity

    Paper | Github

    1. propose hybrid region representation for referring : region name + coordinates + mask pooled feature by Spatial-aware visual sampler
    2. grounding through bbox
    Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

    Paper

    1. propose a bunch of improvements on Ferret v1
    2. including any-resolution (patches) for larger resolution
    3. DINOv2 Encoder for local feature extraction
    4. and High-resolution Dense Alignment stage between SFT and instruction turning.
    u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

    Paper | Github

    1. propose to use different decoder for grounding (SAM for segmentation, Grounding DINO for detection)
    GSVA: Generalized Segmentation via Multimodal Large Language Models

    Paper | Github

    1. propose to Generalized Referring Expression Segmentation (GRES) in grounding LLM
      1. multiple object to ground
      2. need to reject null target
    2. propose to use multple [SEG] token to ground multiple objects (indicted by the texts before the [SEG] token), and [REJ] token to rej null target
    NExT-Chat: An LMM for Chat, Detection and Segmentation

    Paper | Github | Project

    1. propose box encoder-decoder for referring and grounding
    2. for grounding, use token to indicate the presence of a grounding output and input the latent embedding to the box decoder (mask decoder e.g. SAM) for box (mask) generation
    3. for referring, use boxes to represent referred region and use box encoder to encode the referred boxes into features, which is input to LLM.
    4. propose a cycle consistency loss for regularization of box encoder-decoder 图 0
    PerceptionGPT: Effectively Fusing Visual Perception into LLM

    Paper

    1. similar to NExT-Chat, propose box encoder-decoder to encode and decode boxes, but seems to only focus on grounding without referring
    2. One possible intriguing point: grounding output indicator <vis> is used to indicate the presence of grounding output (as usual) but the is replaced by the encoder's output feature in the LLM input. 图 1
    Kosmos-2: Grounding Multimodal Large Language Models to the World

    Paper | Github

    1. build a web-scale grounding dataset by web-scale data (COYO-700M & LAION-2B etc) and vision detector (GLIP)
    2. following pix2seq, divide the image into PxP grids and introduce PxP new tokens to represent
    3. Use <box></box> to represent a bbox, with <delim> to separate multiple boxes (if there are multiple boxes)
    4. Use markdown-like grammar to reference grounded text with <p> </p> e.g. 图 2
    Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

    Paper | Github

    1. propose to use normalized boxes for unified grounding and referring
    2. Use texts to represent all normalized boxes (directly tokenized by text tokenizer) and input to LLM
    Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

    Paper | Github | Project

    1. Propose to ground and refer with a set of proposed regions.
    2. Change a Deformable DETR detection head into binary classifier to propose ROI and use AlignROI pooling to get the region feature
    LISA: Reasoning Segmentation via Large Language Model

    Paper | Github

    1. Introduce the reasoning segmentation task and establish a reasoning segmentation benchmark.
    2. Propose LISA model, which represents the segmentation mask as an embedding and incorporates new segmentation capabilities.
    GLaMM: Pixel Grounding Large Multimodal Model

    Paper | Github | Project

    1. Introduces Grounded Conversation Generation (GCG) task combining phrase grounding, referring expression segmentation, and vision-language conversations.
    2. Proposes a scalable pipeline to curate GranD (Grounding-anything Dataset) with 7.5M unique concepts grounded in 810M regions with segmentation masks, 214k GranDf and a ~1000 evaluation set.
    3. Architecture: Global Image Encoder for holistic understanding, Region Encoder with RoI pooling for regions referring, LLM generates responses with grounding tokens , and use Pixel Decoder (SAM) to decode segmentation masks from tokens' latent

    🔥 Multi-modality

    GroundingGPT:Language Enhanced Multi-modal Grounding Model

    Paper | Github

    1. grounding and referring of multi-modality in text
      1. bounding box by four relative coordinate values:[x1, y1, x2, y2]
      2. video timestamps by two two-digit decimals: {t1, t2}
    2. curate dataset for three stage training 图 1

    williamium3000/awesome-mllm-grounding

    Awesome paper for multi-modal llm with grounding ability

    21

    25 commits

    updated Oct 11, 2025

    See the code

    README

    Awesome-Multimodal-Large-Language-Models-With-Grounding

    A curated list of Multimodal Large Language Models (or Large Vision Language Model) with grounding ability.

    Table of Contents

    🔥 Large Vision-Language Model

    Grounding

    FormatDescPaper
    Decoder on latentleverage a decoder to groundPerceptionGPT, NExT-Chat, PSALM, PixelLM, u-LLaVA, GSVA, ChatterBox, GLaMM
    Output numerical coordinatesdirect output numerical tokensShikra, VisionLLM, Ferret, Ferret2, CogVLM
    Output token coordinatesoutput new tokens added to refer positionsKosmos-2
    Pixel spaceoutput in discrete pixel space encoded by VQGANUnified-IO, Unified-IO 2
    Proposal retrievalretrieval from region candidatesLLM-Seg, Kosmos-2, GROUNDHOG

    Referring

    FormatDescPaper
    PoolingLeverage Mask Pooling / RoI Pooling / RoI Align to obtain features from the im encoder outputGroma, GPT4RoI, Osprey, PSALM, GROUNDHOG, Ferret, Ferret2, PVIT, ChatterBox, GLaMM
    Numerical coordinatesLeverage numerical coordinates for referring (bbox / sampled points in mask)Shikra, PerceptionGPT (w/ encoder), NExT-Chat (w/ encoder), CogVLM
    Token coordinatesAdd new tokens to vocab to present spatial positionsKosmos-2
    • w/ encoder: refers to using a encoder to encode the input coordinates.

    Training Dataset

    DatasetSourceData SourceQuantityCnstruction Method
    GRITFerretCOYO-700M, LAION-2B-
  • Templates are used to convert data.
  • SAM is used to generate masks for free-form referring.
  • ChatGPT4 is used to generate dialogues with bbox.
  • Use GLIPv2 to ground groundable nouns in LLaVA-158k.
  • Negative mining: generate negative yes/or question
  • Shikra-RDShikraFlickr30K Entities5,922 QA pairsChatGPT4 ==> Referential Dialogue (CoT dialogues with grounding & referring)
    CB-300KChatterBoxVG717,075 QA pairs4 subsets.
  • CB-MRG: Use ChatGPT to write dialogues with bbox
  • CB-LC, extend strict relation (from scene graph) to multi-turn QA with ChatGPT
  • CB-REF REG task
  • CB-GND: grounding task
  • GranDGLaMMSA-1B11M images with 7.5M unique concepts and 810M regions.Automated annotation pipeline with SAM for dense pixel-wise grounding. Used for pretraining.
    GranD-fGLaMMGranD (refined), Flickr30K, RefCOCOg, and PSG~214K image-grounded text pairsRefined subset of GranD for fine-tuning, with 1000 images held out for human-annotated evaluation

    Training Recipe

    ModelRecipe
    Ferret
  • Use LLaVA pretrained
  • SFT on GRIT
  • Ferret2
  • image-caption alignment on 1.4M image-text pairs
  • high-resolution dense alignment with template referring & grounding
  • instruction tuning with GRIT, VQA and OCR (VQA and OCR are augmented with GLIPv2 bbox)
  • ChatterBoxTrainable: LoRA and location decoder
  • warm up training with visual grounding only dataset.
  • instruction tuning with CB-300K
  • GPT4RoI
  • Use LLaVA pretrained
  • pretrain region feature extractor with text-region datasets (COCO, RefCOCO, RefCOCO+)
  • train connector, region feature extractor and LLM to follow instructions
  • GLaMM
  • Use GPT4RoI pretrained
  • pretrain on 11M GranD with LoRA
  • finetune on GranD-f, LLaVA-Instruct150K and LLaVA-Instruct-80K
  • Evaluation Dataset

    DatasetSourceData SourceQuantityCnstruction Method
    Ferret BenchFerretCOCO validation set120
  • Referring Description: models are asked to describe a referred region based on its interaction with surrounding objects.
  • Referring Reasoning: models need to reason on top of one or more referred regions correctly.
  • Grounding in Conversation: models are required to reason correctly and accurately ground/localize the objects/regions necessary for the reasoning.
  • Paper List

    GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

    Paper | Github

    1. propose referring for mllm by replacing placeholder <region_i> by feature obtained by mask pooling
    Osprey: Pixel Understanding with Visual Instruction Tuning

    Paper | Github

    1. similar to GPT4RoI, Osprey also use mask representation to refer to entities in images.
    2. It uses mask pooling to extract semantic features from image encoder and combines with a location extractor to process the mask and output spatial token.
    VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

    Paper | Github

    1. unified interface for vision and vl tasks: points for detection, sample points for instance seg ==> instruction format for training
    2. extra tokens & output-format-as-query to decode (faster)
    Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

    Paper | Github | Project

    1. creates a unified IO for all sorts of vision and vl task (into discrete tokens)
    2. using t5-like encoder-decoder arch
    Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

    Paper | Github | Project

    1. following Unified-IO v1, creates a unified IO for all sorts of modalities including image, masks, bboxes, audios (into discrete tokens)
      1. dense masks are all binary, unlike v1 which specifies the color in text instruction (model struggles to follow)
    2. propose 2D Rotary Embedding, QK Normalization and Scaled Cosine Attention to stabilize training and scaling
    3. Mixture of Denoisers taining objectives
    4. instruction tuning of 220 tasks drawn from over 120 external datasets
    PixelLM: Pixel Reasoning with Large Multimodal Model

    Paper | Github | Project

    1. learnable seg tokens + light-weight decoder
    2. a bunch of tricks:
      1. N x L seg tokens for L level multi-scale vision features. N tokens within each group for better modeling
      2. reweighted loss on regions with overlapping predictions
    PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model

    Paper | Github

    1. new paradigm: first generate mask proposal, then genereate mask and classification (following mask2former)
    2. instruction prompt + conditional prompt + candidate masks token
      1. three types of conditional prompt: classes, sentence (ref seg) and visual cues (point, scribbles, boxes, etc)
      2. conditional prompt => condition embed, candidate masks token => mask embed.
      3. condition embed +mask embed + image feature => mask2former decoder => bipartite matching loss + query-based decoding 图 0
    LLM-Seg: Bridging Image Segmentation and Large Language Model Reasoning

    Paper | Github

    1. Use SAM to generate mask candidates, then fomulate the problem as mask selection (mask classification)
    2. promote LLM-Seg40K dataset, by using LLaVA to generate caption, then GPT4 to generate question-answer pair.
    GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

    Paper | Project

    1. disantengle grounding with referring
    2. grounding as mask selection and train a mask2former+ to generate mask candidates
    3. referring by mask pooling on feature
    4. promote 2.5M M3G2 dataset
    DetGPT: Detect What You Need via Reasoning

    Paper | Github | Project

    1. Follow LLaVA to tune VLM and for vqa
    2. Use grouding DINO to ground response generated by VLM to detect the relevantg entities.
    Ferret: Refer and Ground Anything Anywhere at Any Granularity

    Paper | Github

    1. propose hybrid region representation for referring : region name + coordinates + mask pooled feature by Spatial-aware visual sampler
    2. grounding through bbox
    Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

    Paper

    1. propose a bunch of improvements on Ferret v1
    2. including any-resolution (patches) for larger resolution
    3. DINOv2 Encoder for local feature extraction
    4. and High-resolution Dense Alignment stage between SFT and instruction turning.
    u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

    Paper | Github

    1. propose to use different decoder for grounding (SAM for segmentation, Grounding DINO for detection)
    GSVA: Generalized Segmentation via Multimodal Large Language Models

    Paper | Github

    1. propose to Generalized Referring Expression Segmentation (GRES) in grounding LLM
      1. multiple object to ground
      2. need to reject null target
    2. propose to use multple [SEG] token to ground multiple objects (indicted by the texts before the [SEG] token), and [REJ] token to rej null target
    NExT-Chat: An LMM for Chat, Detection and Segmentation

    Paper | Github | Project

    1. propose box encoder-decoder for referring and grounding
    2. for grounding, use token to indicate the presence of a grounding output and input the latent embedding to the box decoder (mask decoder e.g. SAM) for box (mask) generation
    3. for referring, use boxes to represent referred region and use box encoder to encode the referred boxes into features, which is input to LLM.
    4. propose a cycle consistency loss for regularization of box encoder-decoder 图 0
    PerceptionGPT: Effectively Fusing Visual Perception into LLM

    Paper

    1. similar to NExT-Chat, propose box encoder-decoder to encode and decode boxes, but seems to only focus on grounding without referring
    2. One possible intriguing point: grounding output indicator <vis> is used to indicate the presence of grounding output (as usual) but the is replaced by the encoder's output feature in the LLM input. 图 1
    Kosmos-2: Grounding Multimodal Large Language Models to the World

    Paper | Github

    1. build a web-scale grounding dataset by web-scale data (COYO-700M & LAION-2B etc) and vision detector (GLIP)
    2. following pix2seq, divide the image into PxP grids and introduce PxP new tokens to represent
    3. Use <box></box> to represent a bbox, with <delim> to separate multiple boxes (if there are multiple boxes)
    4. Use markdown-like grammar to reference grounded text with <p> </p> e.g. 图 2
    Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

    Paper | Github

    1. propose to use normalized boxes for unified grounding and referring
    2. Use texts to represent all normalized boxes (directly tokenized by text tokenizer) and input to LLM
    Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

    Paper | Github | Project

    1. Propose to ground and refer with a set of proposed regions.
    2. Change a Deformable DETR detection head into binary classifier to propose ROI and use AlignROI pooling to get the region feature
    LISA: Reasoning Segmentation via Large Language Model

    Paper | Github

    1. Introduce the reasoning segmentation task and establish a reasoning segmentation benchmark.
    2. Propose LISA model, which represents the segmentation mask as an embedding and incorporates new segmentation capabilities.
    GLaMM: Pixel Grounding Large Multimodal Model

    Paper | Github | Project

    1. Introduces Grounded Conversation Generation (GCG) task combining phrase grounding, referring expression segmentation, and vision-language conversations.
    2. Proposes a scalable pipeline to curate GranD (Grounding-anything Dataset) with 7.5M unique concepts grounded in 810M regions with segmentation masks, 214k GranDf and a ~1000 evaluation set.
    3. Architecture: Global Image Encoder for holistic understanding, Region Encoder with RoI pooling for regions referring, LLM generates responses with grounding tokens , and use Pixel Decoder (SAM) to decode segmentation masks from tokens' latent

    🔥 Multi-modality

    GroundingGPT:Language Enhanced Multi-modal Grounding Model

    Paper | Github

    1. grounding and referring of multi-modality in text
      1. bounding box by four relative coordinate values:[x1, y1, x2, y2]
      2. video timestamps by two two-digit decimals: {t1, t2}
    2. curate dataset for three stage training 图 1