docling-project/DeskForge-1M

Dataset

DeskForge-1M

7

33 commits

3 linked in READMEs

updated Oct 5, 2026

See the code

README

DeskForge-1M

DeskForge-1M is a corpus of 1.21M annotated desktop screenshots with 159.7M element instances and 917K recorded click transitions, generated with DeskForge, a controllable desktop environment that composes and explores real applications.

Project page · Code · Model · Paper

A. Said Gurbuz · Ahmed Nassar · Sunghwan Hong · Marc Pollefeys · Peter W. J. Staar
ETH Zurich · IBM Research Zurich · Microsoft

DeskForge overview

Every screenshot is annotated densely: not one target, but every visible element, with its type, text, geometry, interaction properties, hierarchy and owning window. Overlapping windows are resolved, so each element records both its full extent and the fragments that are actually visible. Scenes combine several real applications with varied content, window layouts, appearance presets and display resolutions, and short click explorations link each action to the screen before and after it.

At a glance

Observations1,207,368 screenshots in 323,731 scenes
Element instances159.7M (132 per screen on average)
Click transitions917,211, from 135,731 exploration episodes
Instructions663,635 clicks with natural-language instructions
Applications19 real Linux desktop applications, several per screen
Appearance7 presets, from classic Linux to Windows- and macOS-inspired styles
Resolutions7, from 1366×768 to 3840×2160

Splits

Splits are made per scene, so all frames of a scene stay together. Three test splits hold an attribute out of training entirely, which measures generalization to desktop configurations never seen in training.

splitheld outobservationstransitionsinstructions
train999,494760,850551,651
val11,2798,6676,281
test_idnew scenes with seen attributes22,55017,38312,589
test_appGNOME System Monitor, Pluma, Xarchiver43,20632,84822,729
test_themethe Quartz Night Nord preset51,62738,67127,175
test_resolution2880×180079,21258,79243,210

A held-out application is excluded from every scene that contains it, in any frame. See docs/splits.md for how each axis was chosen.

What a sample contains

The default configuration streams the observations as WebDataset shards. Each observation is four members sharing one key:

membercontent
pngthe screenshot
leaf.jsonevery visible element: type, role, name and visible text, rect and visible_fragments, interaction state, parent and reading order, owning application and window
screentag.txtthe screen serialized as ScreenTag markup, on a 0–500 grid normalized to the screenshot
record.jsonscene metadata (applications, preset, resolution, window stack), eligibility flags, and the action that produced this state

Alongside the shards, index/ holds Parquet tables for observations, transitions, instructions, episodes and scenes: index/transitions/<split>.parquet lists every recorded click with its before and after observation keys, target element and effect, and index/instructions/<split>.parquet gives the clicks their natural-language instructions. demo/ holds browsing samples: stratified observations with images inline (downscaled to 1024 px) and the transitions_preview shards below. The field reference is in docs/schema.md.

Browsing transitions

The transitions_preview subset shows recorded clicks in the Dataset Viewer: 100 per split, each from a different episode, varied over applications, element types, appearance presets and resolutions. A row reads as one step:

membercontent
instruction.txta natural-language instruction for the click
before.pngthe screen the click was taken on
target.pngthe clicked element, cropped from the before screen and outlined
after.pngthe screen the click produced
action.txtthe click, its target element and application
transition.jsoninstruction variants, action and target geometry, effect, scene

Screenshots are byte-identical copies of the corpus members. The subset is for browsing; all transitions and instructions are in index/ and the shards.

Instructions

663,635 recorded clicks come with natural-language instructions, synthesized with Qwen3.6-27B from the click, its target and the screens before and after it. An instruction states one coordinate-free goal that can be carried out from the before screen, in a standard and a more detailed_contextual style; primary_instruction is the standard one where it exists, and referring_expression describes the target on the before screen. The 551,651 training instructions are the pool for grounding training, and the 105,703 in the four test splits are the grounding evaluation examples.

Loading

from datasets import load_dataset

# Stream observations from any split.
ds = load_dataset("docling-project/DeskForge-1M", split="test_app", streaming=True)
sample = next(iter(ds))
sample["png"], sample["leaf.json"], sample["screentag.txt"], sample["record.json"]

# Browse recorded clicks: instruction, before, clicked element, after.
preview = load_dataset("docling-project/DeskForge-1M", "transitions_preview", split="val")

# Transitions: before/after keys, the clicked target and what changed.
transitions = load_dataset(
    "parquet",
    data_files="hf://datasets/docling-project/DeskForge-1M/index/transitions/val.parquet",
    split="train",
)

# Instructions: one row per click, joined to transitions by transition_id.
instructions = load_dataset(
    "parquet",
    data_files="hf://datasets/docling-project/DeskForge-1M/index/instructions/val.parquet",
    split="train",
)

For full passes, read the shards with webdataset; the Parquet indexes select subsets before streaming, and give the byte offset of every member for random access:

import webdataset as wds
ds = (wds.WebDataset("data/train/part-{00000..00903}.tar", shardshuffle=True)
        .decode("pil")
        .to_tuple("png", "leaf.json", "screentag.txt", "record.json"))

import pyarrow.parquet as pq
obs = pq.read_table("index/observations/test_id.parquet")
heavily_occluded = obs.filter(obs["occluded_ratio"] > 0.5)

examples/ contains runnable scripts for streaming states, pairing transitions and fetching a single observation by key.

Recommended filters. state_train_eligible selects the still-image view (excluding near-duplicate scenes and frames where a click changed nothing), and transition_train_eligible selects the action view, which keeps no-op clicks as supervision.

Annotation quality

Annotations are generated automatically from the accessibility tree and reconciled with the screenshot and the window stack. Captures that fail the automated audits (element coverage against rendered pixels, blank-widget suppression, window ownership and fragment containment) are not included. A human audit finds 99.8% of sampled element annotations correct and 97.8% of sampled instructions sound.

Limitations

All screenshots come from one Linux backend (Xfce): the Windows- and macOS-inspired presets reproduce the look of those systems rather than run them. Actions are clicks from undirected exploration, not goal-directed demonstrations; instructions are written for single clicks after the fact. Because annotations transcribe what is on screen, text drawn by applications (for example session paths or live web content in browser scenes) is part of the data.

License

The dataset is released under the MIT License. Application interfaces, themes, icons, wallpapers and web content visible in the screenshots remain the property of their respective owners.

Citation

@misc{gurbuz2026deskforge,
      title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
      author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
      year={2026},
      eprint={2610.02320},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02320},
}
accessibility
computer-use
desktop
gui-grounding
screen-parsing
webdataset

docling-project/DeskForge-1M

Dataset

DeskForge-1M

7

33 commits

3 linked in READMEs

updated Oct 5, 2026

See the code

README

DeskForge-1M

DeskForge-1M is a corpus of 1.21M annotated desktop screenshots with 159.7M element instances and 917K recorded click transitions, generated with DeskForge, a controllable desktop environment that composes and explores real applications.

Project page · Code · Model · Paper

A. Said Gurbuz · Ahmed Nassar · Sunghwan Hong · Marc Pollefeys · Peter W. J. Staar
ETH Zurich · IBM Research Zurich · Microsoft

DeskForge overview

Every screenshot is annotated densely: not one target, but every visible element, with its type, text, geometry, interaction properties, hierarchy and owning window. Overlapping windows are resolved, so each element records both its full extent and the fragments that are actually visible. Scenes combine several real applications with varied content, window layouts, appearance presets and display resolutions, and short click explorations link each action to the screen before and after it.

At a glance

Observations1,207,368 screenshots in 323,731 scenes
Element instances159.7M (132 per screen on average)
Click transitions917,211, from 135,731 exploration episodes
Instructions663,635 clicks with natural-language instructions
Applications19 real Linux desktop applications, several per screen
Appearance7 presets, from classic Linux to Windows- and macOS-inspired styles
Resolutions7, from 1366×768 to 3840×2160

Splits

Splits are made per scene, so all frames of a scene stay together. Three test splits hold an attribute out of training entirely, which measures generalization to desktop configurations never seen in training.

splitheld outobservationstransitionsinstructions
train999,494760,850551,651
val11,2798,6676,281
test_idnew scenes with seen attributes22,55017,38312,589
test_appGNOME System Monitor, Pluma, Xarchiver43,20632,84822,729
test_themethe Quartz Night Nord preset51,62738,67127,175
test_resolution2880×180079,21258,79243,210

A held-out application is excluded from every scene that contains it, in any frame. See docs/splits.md for how each axis was chosen.

What a sample contains

The default configuration streams the observations as WebDataset shards. Each observation is four members sharing one key:

membercontent
pngthe screenshot
leaf.jsonevery visible element: type, role, name and visible text, rect and visible_fragments, interaction state, parent and reading order, owning application and window
screentag.txtthe screen serialized as ScreenTag markup, on a 0–500 grid normalized to the screenshot
record.jsonscene metadata (applications, preset, resolution, window stack), eligibility flags, and the action that produced this state

Alongside the shards, index/ holds Parquet tables for observations, transitions, instructions, episodes and scenes: index/transitions/<split>.parquet lists every recorded click with its before and after observation keys, target element and effect, and index/instructions/<split>.parquet gives the clicks their natural-language instructions. demo/ holds browsing samples: stratified observations with images inline (downscaled to 1024 px) and the transitions_preview shards below. The field reference is in docs/schema.md.

Browsing transitions

The transitions_preview subset shows recorded clicks in the Dataset Viewer: 100 per split, each from a different episode, varied over applications, element types, appearance presets and resolutions. A row reads as one step:

membercontent
instruction.txta natural-language instruction for the click
before.pngthe screen the click was taken on
target.pngthe clicked element, cropped from the before screen and outlined
after.pngthe screen the click produced
action.txtthe click, its target element and application
transition.jsoninstruction variants, action and target geometry, effect, scene

Screenshots are byte-identical copies of the corpus members. The subset is for browsing; all transitions and instructions are in index/ and the shards.

Instructions

663,635 recorded clicks come with natural-language instructions, synthesized with Qwen3.6-27B from the click, its target and the screens before and after it. An instruction states one coordinate-free goal that can be carried out from the before screen, in a standard and a more detailed_contextual style; primary_instruction is the standard one where it exists, and referring_expression describes the target on the before screen. The 551,651 training instructions are the pool for grounding training, and the 105,703 in the four test splits are the grounding evaluation examples.

Loading

from datasets import load_dataset

# Stream observations from any split.
ds = load_dataset("docling-project/DeskForge-1M", split="test_app", streaming=True)
sample = next(iter(ds))
sample["png"], sample["leaf.json"], sample["screentag.txt"], sample["record.json"]

# Browse recorded clicks: instruction, before, clicked element, after.
preview = load_dataset("docling-project/DeskForge-1M", "transitions_preview", split="val")

# Transitions: before/after keys, the clicked target and what changed.
transitions = load_dataset(
    "parquet",
    data_files="hf://datasets/docling-project/DeskForge-1M/index/transitions/val.parquet",
    split="train",
)

# Instructions: one row per click, joined to transitions by transition_id.
instructions = load_dataset(
    "parquet",
    data_files="hf://datasets/docling-project/DeskForge-1M/index/instructions/val.parquet",
    split="train",
)

For full passes, read the shards with webdataset; the Parquet indexes select subsets before streaming, and give the byte offset of every member for random access:

import webdataset as wds
ds = (wds.WebDataset("data/train/part-{00000..00903}.tar", shardshuffle=True)
        .decode("pil")
        .to_tuple("png", "leaf.json", "screentag.txt", "record.json"))

import pyarrow.parquet as pq
obs = pq.read_table("index/observations/test_id.parquet")
heavily_occluded = obs.filter(obs["occluded_ratio"] > 0.5)

examples/ contains runnable scripts for streaming states, pairing transitions and fetching a single observation by key.

Recommended filters. state_train_eligible selects the still-image view (excluding near-duplicate scenes and frames where a click changed nothing), and transition_train_eligible selects the action view, which keeps no-op clicks as supervision.

Annotation quality

Annotations are generated automatically from the accessibility tree and reconciled with the screenshot and the window stack. Captures that fail the automated audits (element coverage against rendered pixels, blank-widget suppression, window ownership and fragment containment) are not included. A human audit finds 99.8% of sampled element annotations correct and 97.8% of sampled instructions sound.

Limitations

All screenshots come from one Linux backend (Xfce): the Windows- and macOS-inspired presets reproduce the look of those systems rather than run them. Actions are clicks from undirected exploration, not goal-directed demonstrations; instructions are written for single clicks after the fact. Because annotations transcribe what is on screen, text drawn by applications (for example session paths or live web content in browser scenes) is part of the data.

License

The dataset is released under the MIT License. Application interfaces, themes, icons, wallpapers and web content visible in the screenshots remain the property of their respective owners.

Citation

@misc{gurbuz2026deskforge,
      title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
      author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
      year={2026},
      eprint={2610.02320},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02320},
}
accessibility
computer-use
desktop
gui-grounding
screen-parsing
webdataset