DeskForge-1M is a corpus of 1.21M annotated desktop screenshots with 159.7M element instances and 917K recorded click transitions, generated with DeskForge, a controllable desktop environment that composes and explores real applications.
Project page · Code · Model · Paper
A. Said Gurbuz · Ahmed Nassar · Sunghwan Hong · Marc Pollefeys · Peter W. J. Staar
ETH Zurich · IBM Research Zurich · Microsoft

Every screenshot is annotated densely: not one target, but every visible element, with its type, text, geometry, interaction properties, hierarchy and owning window. Overlapping windows are resolved, so each element records both its full extent and the fragments that are actually visible. Scenes combine several real applications with varied content, window layouts, appearance presets and display resolutions, and short click explorations link each action to the screen before and after it.
| Observations | 1,207,368 screenshots in 323,731 scenes |
| Element instances | 159.7M (132 per screen on average) |
| Click transitions | 917,211, from 135,731 exploration episodes |
| Instructions | 663,635 clicks with natural-language instructions |
| Applications | 19 real Linux desktop applications, several per screen |
| Appearance | 7 presets, from classic Linux to Windows- and macOS-inspired styles |
| Resolutions | 7, from 1366×768 to 3840×2160 |
Splits are made per scene, so all frames of a scene stay together. Three test splits hold an attribute out of training entirely, which measures generalization to desktop configurations never seen in training.
| split | held out | observations | transitions | instructions |
|---|---|---|---|---|
train | 999,494 | 760,850 | 551,651 | |
val | 11,279 | 8,667 | 6,281 | |
test_id | new scenes with seen attributes | 22,550 | 17,383 | 12,589 |
test_app | GNOME System Monitor, Pluma, Xarchiver | 43,206 | 32,848 | 22,729 |
test_theme | the Quartz Night Nord preset | 51,627 | 38,671 | 27,175 |
test_resolution | 2880×1800 | 79,212 | 58,792 | 43,210 |
A held-out application is excluded from every scene that contains it, in any
frame. See docs/splits.md for how each axis was chosen.
The default configuration streams the observations as WebDataset shards. Each observation is four members sharing one key:
| member | content |
|---|---|
png | the screenshot |
leaf.json | every visible element: type, role, name and visible text, rect and visible_fragments, interaction state, parent and reading order, owning application and window |
screentag.txt | the screen serialized as ScreenTag markup, on a 0–500 grid normalized to the screenshot |
record.json | scene metadata (applications, preset, resolution, window stack), eligibility flags, and the action that produced this state |
Alongside the shards, index/ holds Parquet tables for observations,
transitions, instructions, episodes and scenes:
index/transitions/<split>.parquet lists every recorded click with its before
and after observation keys, target element and effect, and
index/instructions/<split>.parquet gives the clicks their natural-language
instructions. demo/ holds browsing samples: stratified observations with
images inline (downscaled to 1024 px) and the transitions_preview shards
below. The field reference is in docs/schema.md.
The transitions_preview subset shows recorded clicks in the Dataset Viewer:
100 per split, each from a different episode, varied over applications, element
types, appearance presets and resolutions. A row reads as one step:
| member | content |
|---|---|
instruction.txt | a natural-language instruction for the click |
before.png | the screen the click was taken on |
target.png | the clicked element, cropped from the before screen and outlined |
after.png | the screen the click produced |
action.txt | the click, its target element and application |
transition.json | instruction variants, action and target geometry, effect, scene |
Screenshots are byte-identical copies of the corpus members. The subset is for
browsing; all transitions and instructions are in index/ and the shards.
663,635 recorded clicks come with natural-language instructions, synthesized
with Qwen3.6-27B from the click, its target and the screens before and after
it. An instruction states one coordinate-free goal that can be carried out from
the before screen, in a standard and a more detailed_contextual style;
primary_instruction is the standard one where it exists, and
referring_expression describes the target on the before screen. The 551,651
training instructions are the pool for grounding training, and the 105,703 in
the four test splits are the grounding evaluation examples.
from datasets import load_dataset
# Stream observations from any split.
ds = load_dataset("docling-project/DeskForge-1M", split="test_app", streaming=True)
sample = next(iter(ds))
sample["png"], sample["leaf.json"], sample["screentag.txt"], sample["record.json"]
# Browse recorded clicks: instruction, before, clicked element, after.
preview = load_dataset("docling-project/DeskForge-1M", "transitions_preview", split="val")
# Transitions: before/after keys, the clicked target and what changed.
transitions = load_dataset(
"parquet",
data_files="hf://datasets/docling-project/DeskForge-1M/index/transitions/val.parquet",
split="train",
)
# Instructions: one row per click, joined to transitions by transition_id.
instructions = load_dataset(
"parquet",
data_files="hf://datasets/docling-project/DeskForge-1M/index/instructions/val.parquet",
split="train",
)
For full passes, read the shards with webdataset;
the Parquet indexes select subsets before streaming, and give the byte offset of
every member for random access:
import webdataset as wds
ds = (wds.WebDataset("data/train/part-{00000..00903}.tar", shardshuffle=True)
.decode("pil")
.to_tuple("png", "leaf.json", "screentag.txt", "record.json"))
import pyarrow.parquet as pq
obs = pq.read_table("index/observations/test_id.parquet")
heavily_occluded = obs.filter(obs["occluded_ratio"] > 0.5)
examples/ contains runnable scripts for streaming states, pairing
transitions and fetching a single observation by key.
Recommended filters. state_train_eligible selects the still-image view
(excluding near-duplicate scenes and frames where a click changed nothing), and
transition_train_eligible selects the action view, which keeps no-op clicks as
supervision.
Annotations are generated automatically from the accessibility tree and reconciled with the screenshot and the window stack. Captures that fail the automated audits (element coverage against rendered pixels, blank-widget suppression, window ownership and fragment containment) are not included. A human audit finds 99.8% of sampled element annotations correct and 97.8% of sampled instructions sound.
All screenshots come from one Linux backend (Xfce): the Windows- and macOS-inspired presets reproduce the look of those systems rather than run them. Actions are clicks from undirected exploration, not goal-directed demonstrations; instructions are written for single clicks after the fact. Because annotations transcribe what is on screen, text drawn by applications (for example session paths or live web content in browser scenes) is part of the data.
The dataset is released under the MIT License. Application interfaces, themes, icons, wallpapers and web content visible in the screenshots remain the property of their respective owners.
@misc{gurbuz2026deskforge,
title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
year={2026},
eprint={2610.02320},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02320},
}
DeskForge-1M is a corpus of 1.21M annotated desktop screenshots with 159.7M element instances and 917K recorded click transitions, generated with DeskForge, a controllable desktop environment that composes and explores real applications.
Project page · Code · Model · Paper
A. Said Gurbuz · Ahmed Nassar · Sunghwan Hong · Marc Pollefeys · Peter W. J. Staar
ETH Zurich · IBM Research Zurich · Microsoft

Every screenshot is annotated densely: not one target, but every visible element, with its type, text, geometry, interaction properties, hierarchy and owning window. Overlapping windows are resolved, so each element records both its full extent and the fragments that are actually visible. Scenes combine several real applications with varied content, window layouts, appearance presets and display resolutions, and short click explorations link each action to the screen before and after it.
| Observations | 1,207,368 screenshots in 323,731 scenes |
| Element instances | 159.7M (132 per screen on average) |
| Click transitions | 917,211, from 135,731 exploration episodes |
| Instructions | 663,635 clicks with natural-language instructions |
| Applications | 19 real Linux desktop applications, several per screen |
| Appearance | 7 presets, from classic Linux to Windows- and macOS-inspired styles |
| Resolutions | 7, from 1366×768 to 3840×2160 |
Splits are made per scene, so all frames of a scene stay together. Three test splits hold an attribute out of training entirely, which measures generalization to desktop configurations never seen in training.
| split | held out | observations | transitions | instructions |
|---|---|---|---|---|
train | 999,494 | 760,850 | 551,651 | |
val | 11,279 | 8,667 | 6,281 | |
test_id | new scenes with seen attributes | 22,550 | 17,383 | 12,589 |
test_app | GNOME System Monitor, Pluma, Xarchiver | 43,206 | 32,848 | 22,729 |
test_theme | the Quartz Night Nord preset | 51,627 | 38,671 | 27,175 |
test_resolution | 2880×1800 | 79,212 | 58,792 | 43,210 |
A held-out application is excluded from every scene that contains it, in any
frame. See docs/splits.md for how each axis was chosen.
The default configuration streams the observations as WebDataset shards. Each observation is four members sharing one key:
| member | content |
|---|---|
png | the screenshot |
leaf.json | every visible element: type, role, name and visible text, rect and visible_fragments, interaction state, parent and reading order, owning application and window |
screentag.txt | the screen serialized as ScreenTag markup, on a 0–500 grid normalized to the screenshot |
record.json | scene metadata (applications, preset, resolution, window stack), eligibility flags, and the action that produced this state |
Alongside the shards, index/ holds Parquet tables for observations,
transitions, instructions, episodes and scenes:
index/transitions/<split>.parquet lists every recorded click with its before
and after observation keys, target element and effect, and
index/instructions/<split>.parquet gives the clicks their natural-language
instructions. demo/ holds browsing samples: stratified observations with
images inline (downscaled to 1024 px) and the transitions_preview shards
below. The field reference is in docs/schema.md.
The transitions_preview subset shows recorded clicks in the Dataset Viewer:
100 per split, each from a different episode, varied over applications, element
types, appearance presets and resolutions. A row reads as one step:
| member | content |
|---|---|
instruction.txt | a natural-language instruction for the click |
before.png | the screen the click was taken on |
target.png | the clicked element, cropped from the before screen and outlined |
after.png | the screen the click produced |
action.txt | the click, its target element and application |
transition.json | instruction variants, action and target geometry, effect, scene |
Screenshots are byte-identical copies of the corpus members. The subset is for
browsing; all transitions and instructions are in index/ and the shards.
663,635 recorded clicks come with natural-language instructions, synthesized
with Qwen3.6-27B from the click, its target and the screens before and after
it. An instruction states one coordinate-free goal that can be carried out from
the before screen, in a standard and a more detailed_contextual style;
primary_instruction is the standard one where it exists, and
referring_expression describes the target on the before screen. The 551,651
training instructions are the pool for grounding training, and the 105,703 in
the four test splits are the grounding evaluation examples.
from datasets import load_dataset
# Stream observations from any split.
ds = load_dataset("docling-project/DeskForge-1M", split="test_app", streaming=True)
sample = next(iter(ds))
sample["png"], sample["leaf.json"], sample["screentag.txt"], sample["record.json"]
# Browse recorded clicks: instruction, before, clicked element, after.
preview = load_dataset("docling-project/DeskForge-1M", "transitions_preview", split="val")
# Transitions: before/after keys, the clicked target and what changed.
transitions = load_dataset(
"parquet",
data_files="hf://datasets/docling-project/DeskForge-1M/index/transitions/val.parquet",
split="train",
)
# Instructions: one row per click, joined to transitions by transition_id.
instructions = load_dataset(
"parquet",
data_files="hf://datasets/docling-project/DeskForge-1M/index/instructions/val.parquet",
split="train",
)
For full passes, read the shards with webdataset;
the Parquet indexes select subsets before streaming, and give the byte offset of
every member for random access:
import webdataset as wds
ds = (wds.WebDataset("data/train/part-{00000..00903}.tar", shardshuffle=True)
.decode("pil")
.to_tuple("png", "leaf.json", "screentag.txt", "record.json"))
import pyarrow.parquet as pq
obs = pq.read_table("index/observations/test_id.parquet")
heavily_occluded = obs.filter(obs["occluded_ratio"] > 0.5)
examples/ contains runnable scripts for streaming states, pairing
transitions and fetching a single observation by key.
Recommended filters. state_train_eligible selects the still-image view
(excluding near-duplicate scenes and frames where a click changed nothing), and
transition_train_eligible selects the action view, which keeps no-op clicks as
supervision.
Annotations are generated automatically from the accessibility tree and reconciled with the screenshot and the window stack. Captures that fail the automated audits (element coverage against rendered pixels, blank-widget suppression, window ownership and fragment containment) are not included. A human audit finds 99.8% of sampled element annotations correct and 97.8% of sampled instructions sound.
All screenshots come from one Linux backend (Xfce): the Windows- and macOS-inspired presets reproduce the look of those systems rather than run them. Actions are clicks from undirected exploration, not goal-directed demonstrations; instructions are written for single clicks after the fact. Because annotations transcribe what is on screen, text drawn by applications (for example session paths or live web content in browser scenes) is part of the data.
The dataset is released under the MIT License. Application interfaces, themes, icons, wallpapers and web content visible in the screenshots remain the property of their respective owners.
@misc{gurbuz2026deskforge,
title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
year={2026},
eprint={2610.02320},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02320},
}