For desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
5
19 commits
3 linked in READMEs
updated Jul 8, 2025
For desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
To further improve data quality, we apply a lightweight cleaning strategy:
This helps ensure that training data remains consistent with actual visual targets, reducing noise from misaligned annotations. While this method may occasionally filter out a small number of false positives, we find such cases account for less than 3% of the data. Refer to our code for details.
Examples from the Aria-UI dataset collection. The blue bounding box shows the derived annotation, while the red bounding boxes are detected using OmniParser. A large green arrow is used to draw attention to the misaligned blue bounding box. Our lightweight cleaning strategy filters out such cases where the annotation does not match the actual UI element.
For desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
5
19 commits
3 linked in READMEs
updated Jul 8, 2025
For desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
To further improve data quality, we apply a lightweight cleaning strategy:
This helps ensure that training data remains consistent with actual visual targets, reducing noise from misaligned annotations. While this method may occasionally filter out a small number of false positives, we find such cases account for less than 3% of the data. Refer to our code for details.
Examples from the Aria-UI dataset collection. The blue bounding box shows the derived annotation, while the red bounding boxes are detected using OmniParser. A large green arrow is used to draw attention to the misaligned blue bounding box. Our lightweight cleaning strategy filters out such cases where the annotation does not match the actual UI element.