37
stars
65
commits
2
linked in READMEs
Mar 10, 2025
updated
Github | arXiv | HF Paper | Spaces | Datasets | Quick Start
ShowUI-desktop-8K is a UI-grounding dataset focused on PC-based grounding, with screenshots and annotations originally sourced from OmniAct.
We utilize GPT-4o to augment the original annotations, enriching them with diverse attributes such as appearance, spatial relationships, and intended functionality.
You can use our rewrite strategy code to augment your own data.

If you find our work helpful, please consider citing our paper.
@misc{lin2024showui,
title={ShowUI: One Vision-Language-Action Model for GUI Visual Agent},
author={Kevin Qinghong Lin and Linjie Li and Difei Gao and Zhengyuan Yang and Shiwei Wu and Zechen Bai and Weixian Lei and Lijuan Wang and Mike Zheng Shou},
year={2024},
eprint={2411.17465},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2411.17465},
}
65 commits
37
stars
65
commits
2
linked in READMEs
Mar 10, 2025
updated
Github | arXiv | HF Paper | Spaces | Datasets | Quick Start
ShowUI-desktop-8K is a UI-grounding dataset focused on PC-based grounding, with screenshots and annotations originally sourced from OmniAct.
We utilize GPT-4o to augment the original annotations, enriching them with diverse attributes such as appearance, spatial relationships, and intended functionality.
You can use our rewrite strategy code to augment your own data.

If you find our work helpful, please consider citing our paper.
@misc{lin2024showui,
title={ShowUI: One Vision-Language-Action Model for GUI Visual Agent},
author={Kevin Qinghong Lin and Linjie Li and Difei Gao and Zhengyuan Yang and Shiwei Wu and Zechen Bai and Weixian Lei and Lijuan Wang and Mike Zheng Shou},
year={2024},
eprint={2411.17465},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2411.17465},
}
65 commits