osunlp/UGround

Model

26

stars

59

commits

6

repos using this model

4

linked in READMEs

Apr 16, 2025

updated

image-text-to-text
llava_llama
safetensors

README

UGround (The Initial LLaVA-based Version)

Update: We have trained stronger models based on Qwen2-VL with the same data. We suggest using them instead for better performance and more convenient training, inference and deployment.

UGround is a strong GUI visual grounding model trained with a simple recipe. Check our homepage and paper for more details. This work is a collaboration between OSU NLP Group and Orby AI. radar

Models

Release Plan

Main Results

GUI Visual Grounding: ScreenSpot (Standard Setting)

Grounding ModelArchSFT dataMobile-TextMobile-IconDesktop-TextDesktop-IconWeb-TextWeb-IconAvg
GPT-422.624.520.211.89.28.816.2
GPT-4o20.224.921.123.612.27.818.3
MiniGPT-v2MiniGPT-v28.46.66.22.96.53.45.7
GromaGroma10.32.64.64.35.73.45.2
FuyuFuyu41.01.333.03.633.94.419.5
Qwen-VLQwen-VL9.54.85.75.03.52.45.2
SeeClickQwen-VLSeeClick78.052.072.230.055.732.553.4
Qwen-GUIQwen-VLGUICourse52.410.945.95.743.013.628.6
UGround-V1LLaVA-UGround-V1UGround-V182.860.382.563.680.470.473.3
Qwen2-VLQwen2-VL61.339.352.045.033.021.842.1
Auguvis-G-7BQwen2-VLAguvis-Stage-188.378.288.170.785.774.881.0
Auguvis-7BQwen2-VLAguvis-Stage-1&295.677.793.867.188.375.283.0
OS-Atlas-Base-4BInternVLOS-Atlas85.758.572.245.782.663.168.0
OS-Atlas-Base-7BQwen2-VLOS-Atlas93.072.991.862.990.974.381.0
ShowUI-GShowUIShowUI91.669.081.859.083.065.575.0
ShowUIShowUIShowUI92.375.576.361.181.763.675.1
IrisIrisSeeClick85.364.286.757.582.671.274.6
Aria-UIAriaAria-UI92.373.893.364.386.576.281.1
UGround-V1-2B (Qwen2-VL)Qwen2-VLUGround-V189.472.088.765.781.368.977.7
UGround-V1-7B (Qwen2-VL)Qwen2-VLUGround-V193.079.993.876.490.984.086.3

GUI Visual Grounding: ScreenSpot (Agent Setting)

PlannerGrounding ModelArchSFT dataMobile-TextMobile-IconDesktop-TextDesktop-IconWeb-TextWeb-IconAvg
GPT-4oQwen-VLQwen-VL21.321.418.610.79.15.814.5
GPT-4oSeeClickQwen-VLSeeClick81.059.869.633.643.926.252.4
GPT-4oQwen-GUIQwen-VLGUICourse67.824.553.116.450.418.538.5
GPT-4oUGround-V1LLaVA-UGround-V1UGround-V193.476.992.867.988.768.981.4
GPT-4oOS-Atlas-Base-4BInternVLOS-Atlas94.173.877.847.186.565.374.1
GPT-4oOS-Atlas-Base-7BQwen2-VLOS-Atlas93.879.990.266.492.679.183.7
GPT-4oUGround-V1-2B (Qwen2-VL)Qwen2-VLUGround-V194.177.792.863.690.070.981.5
GPT-4oUGround-V1-7B (Qwen2-VL)Qwen2-VLUGround-V194.179.993.373.689.673.384.0

image/png

Citation Information

If you find this work useful, please consider citing our papers:

@article{gou2024uground,
        title={Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents},
        author={Boyu Gou and Ruohan Wang and Boyuan Zheng and Yanan Xie and Cheng Chang and Yiheng Shu and Huan Sun and Yu Su},
        journal={arXiv preprint arXiv:2410.05243},
        year={2024},
        url={https://arxiv.org/abs/2410.05243},
      }

@article{zheng2023seeact,
        title={GPT-4V(ision) is a Generalist Web Agent, if Grounded},
        author={Boyuan Zheng and Boyu Gou and Jihyung Kil and Huan Sun and Yu Su},
        journal={arXiv preprint arXiv:2401.01614},
        year={2024},
      }

Contributors

BoyuNLP

55 commits

nielsr

2 commits

EthanKe

1 commits

mznyu

1 commits

osunlp/UGround

Model

26

stars

59

commits

6

repos using this model

4

linked in READMEs

Apr 16, 2025

updated

image-text-to-text
llava_llama
safetensors

README

UGround (The Initial LLaVA-based Version)

Update: We have trained stronger models based on Qwen2-VL with the same data. We suggest using them instead for better performance and more convenient training, inference and deployment.

UGround is a strong GUI visual grounding model trained with a simple recipe. Check our homepage and paper for more details. This work is a collaboration between OSU NLP Group and Orby AI. radar

Models

Release Plan

Main Results

GUI Visual Grounding: ScreenSpot (Standard Setting)

Grounding ModelArchSFT dataMobile-TextMobile-IconDesktop-TextDesktop-IconWeb-TextWeb-IconAvg
GPT-422.624.520.211.89.28.816.2
GPT-4o20.224.921.123.612.27.818.3
MiniGPT-v2MiniGPT-v28.46.66.22.96.53.45.7
GromaGroma10.32.64.64.35.73.45.2
FuyuFuyu41.01.333.03.633.94.419.5
Qwen-VLQwen-VL9.54.85.75.03.52.45.2
SeeClickQwen-VLSeeClick78.052.072.230.055.732.553.4
Qwen-GUIQwen-VLGUICourse52.410.945.95.743.013.628.6
UGround-V1LLaVA-UGround-V1UGround-V182.860.382.563.680.470.473.3
Qwen2-VLQwen2-VL61.339.352.045.033.021.842.1
Auguvis-G-7BQwen2-VLAguvis-Stage-188.378.288.170.785.774.881.0
Auguvis-7BQwen2-VLAguvis-Stage-1&295.677.793.867.188.375.283.0
OS-Atlas-Base-4BInternVLOS-Atlas85.758.572.245.782.663.168.0
OS-Atlas-Base-7BQwen2-VLOS-Atlas93.072.991.862.990.974.381.0
ShowUI-GShowUIShowUI91.669.081.859.083.065.575.0
ShowUIShowUIShowUI92.375.576.361.181.763.675.1
IrisIrisSeeClick85.364.286.757.582.671.274.6
Aria-UIAriaAria-UI92.373.893.364.386.576.281.1
UGround-V1-2B (Qwen2-VL)Qwen2-VLUGround-V189.472.088.765.781.368.977.7
UGround-V1-7B (Qwen2-VL)Qwen2-VLUGround-V193.079.993.876.490.984.086.3

GUI Visual Grounding: ScreenSpot (Agent Setting)

PlannerGrounding ModelArchSFT dataMobile-TextMobile-IconDesktop-TextDesktop-IconWeb-TextWeb-IconAvg
GPT-4oQwen-VLQwen-VL21.321.418.610.79.15.814.5
GPT-4oSeeClickQwen-VLSeeClick81.059.869.633.643.926.252.4
GPT-4oQwen-GUIQwen-VLGUICourse67.824.553.116.450.418.538.5
GPT-4oUGround-V1LLaVA-UGround-V1UGround-V193.476.992.867.988.768.981.4
GPT-4oOS-Atlas-Base-4BInternVLOS-Atlas94.173.877.847.186.565.374.1
GPT-4oOS-Atlas-Base-7BQwen2-VLOS-Atlas93.879.990.266.492.679.183.7
GPT-4oUGround-V1-2B (Qwen2-VL)Qwen2-VLUGround-V194.177.792.863.690.070.981.5
GPT-4oUGround-V1-7B (Qwen2-VL)Qwen2-VLUGround-V194.179.993.373.689.673.384.0

image/png

Citation Information

If you find this work useful, please consider citing our papers:

@article{gou2024uground,
        title={Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents},
        author={Boyu Gou and Ruohan Wang and Boyuan Zheng and Yanan Xie and Cheng Chang and Yiheng Shu and Huan Sun and Yu Su},
        journal={arXiv preprint arXiv:2410.05243},
        year={2024},
        url={https://arxiv.org/abs/2410.05243},
      }

@article{zheng2023seeact,
        title={GPT-4V(ision) is a Generalist Web Agent, if Grounded},
        author={Boyuan Zheng and Boyu Gou and Jihyung Kil and Huan Sun and Yu Su},
        journal={arXiv preprint arXiv:2401.01614},
        year={2024},
      }

Contributors

BoyuNLP

55 commits

nielsr

2 commits

EthanKe

1 commits

mznyu

1 commits