DeskForge-Qwen3.5-4B is Qwen3.5-4B
fine-tuned on 200K grounding examples from
DeskForge-1M.
Given a desktop screenshot and a natural-language instruction, it answers with
the next action as a pyautogui call whose coordinates are fractions of the
screen.
Project page · Dataset · Code · Paper
Grounding on DeskForge-1M held-out conditions (accuracy, %). New Scenes are unseen scenes with seen attributes; Theme, App and Resolution hold an appearance preset, three applications and a display resolution out of training entirely.
| Model | New Scenes | Theme | App | Resolution | Mean |
|---|---|---|---|---|---|
| UI-TARS-1.5-7B | 67.68 | 64.96 | 66.68 | 61.67 | 65.25 |
| GroundNext-7B | 73.57 | 67.67 | 72.10 | 68.87 | 70.55 |
| UI-Venus-2-9B | 81.45 | 77.44 | 79.40 | 77.40 | 78.92 |
| Qwen3.5-4B | 78.56 | 74.03 | 76.11 | 76.34 | 76.26 |
| DeskForge-Qwen3.5-4B | 90.14 | 87.36 | 87.50 | 85.15 | 87.54 |
Grounding on external GUI benchmarks (accuracy, %), none of which is used for training:
| Model | ScreenSpot-Pro | ScreenSpot-v2 | OSWorld-G | UI-Vision | MMBench-GUI |
|---|---|---|---|---|---|
| Qwen3.5-4B | 30.17 | 80.11 | 51.24 | 17.24 | 61.83 |
| DeskForge-Qwen3.5-4B | 41.68 | 90.33 | 61.35 | 28.77 | 76.82 |
Long-horizon task completion (tasks solved). A fixed Qwen3.6-27B planner decides each step and the action model locates its targets, so only the action model differs between the rows:
| Action model | WebArena-Infinity (119 tasks) | OpenApps (100 tasks) |
|---|---|---|
| Qwen3.5-4B | 31 | 3 |
| DeskForge-Qwen3.5-4B | 50 | 15 |
The model takes the system prompt it was trained with, one screenshot, and the task.
import re
from PIL import Image
MODEL = "docling-project/DeskForge-Qwen3.5-4B"
SYSTEM_PROMPT = """You are a computer-use agent operating a desktop graphical interface. At each step you see the user's task, a screenshot of the current screen, and the actions you have already taken. Reply with the next action as pyautogui code and nothing else -- no explanation, no code fence, no commentary.
Coordinates are fractions of the screen, not pixels: x runs from 0.0 at the left edge to 1.0 at the right edge, y from 0.0 at the top to 1.0 at the bottom. Write both with four decimals.
These are the only actions available:
pyautogui.click(x=0.0000, y=0.0000)
pyautogui.doubleClick(x=0.0000, y=0.0000)
pyautogui.rightClick(x=0.0000, y=0.0000)
pyautogui.middleClick(x=0.0000, y=0.0000)
computer.tripleClick(x=0.0000, y=0.0000)
pyautogui.moveTo(x=0.0000, y=0.0000)
pyautogui.dragTo(x=0.0000, y=0.0000, button='left')
pyautogui.scroll(-4)
pyautogui.hscroll(4)
pyautogui.write(message='text to type')
pyautogui.press('enter')
pyautogui.hotkey(['ctrl', 'c'])
computer.wait()
computer.terminate(status='success')
To scroll at a particular place, move there first and then scroll. When the task is finished, or cannot be finished, end with computer.terminate."""
def fit(image, max_pixels=2_097_152):
"""Downscale to at most max_pixels (even sides, Lanczos), as in training."""
w, h = image.size
if w * h <= max_pixels:
return image
s = (max_pixels / (w * h)) ** 0.5
return image.resize((max(2, int(w * s) // 2 * 2), max(2, int(h * s) // 2 * 2)), Image.LANCZOS)
def user_text(instruction):
return f"Task: {instruction}\n\nActions already taken:\n(none -- this is the first step)\n\nNext action:"
def chat(instruction):
return [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": user_text(instruction)}]},
]
def to_pixels(action, screenshot):
x, y = map(float, re.search(r"x=([\d.]+), y=([\d.]+)", action).groups())
return round(x * screenshot.width), round(y * screenshot.height)
screenshot = Image.open("screenshot.png").convert("RGB")
instruction = "Open the File menu of the text editor."
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(MODEL, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(MODEL)
prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
add_generation_prompt=True, enable_thinking=False)
inputs = processor(text=[prompt], images=[fit(screenshot)], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
action = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()
print(action, to_pixels(action, screenshot)) # e.g. pyautogui.click(x=0.0412, y=0.0535) (105, 77)
from transformers import AutoProcessor
from vllm import LLM, SamplingParams
processor = AutoProcessor.from_pretrained(MODEL)
llm = LLM(model=MODEL, max_model_len=16384, limit_mm_per_prompt={"image": 1})
prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
add_generation_prompt=True, enable_thinking=False)
outputs = llm.generate({"prompt": prompt, "multi_modal_data": {"image": fit(screenshot)}},
SamplingParams(temperature=0.0, max_tokens=128))
action = outputs[0].outputs[0].text.strip()
print(action, to_pixels(action, screenshot))
vllm serve docling-project/DeskForge-Qwen3.5-4B --max-model-len 16384
import base64
import io
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
buffer = io.BytesIO()
fit(screenshot).save(buffer, format="PNG")
image_url = "data:image/png;base64," + base64.b64encode(buffer.getvalue()).decode()
response = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": image_url}},
{"type": "text", "text": user_text(instruction)},
]},
],
max_tokens=128,
temperature=0.0,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
action = response.choices[0].message.content.strip()
print(action, to_pixels(action, screenshot))
| Base model | Qwen/Qwen3.5-4B |
| Data | 200,000 single-target grounding examples from the DeskForge-1M training split: a screenshot, an instruction, and a click on the target |
| Method | LoRA on the language model (rank 8, α 32, dropout 0.05), vision encoder and aligner frozen, merged into the base weights |
| Optimization | One epoch, 3,125 updates, effective batch size 64, AdamW, learning rate 10⁻⁴ with cosine decay and 3% warm-up, BF16 |
| Inputs | Screenshots capped at 2,097,152 pixels; sequences up to 4,096 tokens |
| Hardware | 8 × NVIDIA H100 80GB |
The model predicts the next GUI action, mainly where to click, from a screenshot and an instruction, and serves as the action model under a planner in multi-step tasks.
Apache 2.0, following the base model; see LICENSE.
@misc{gurbuz2026deskforge,
title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
year={2026},
eprint={2610.02320},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02320},
}
DeskForge-Qwen3.5-4B is Qwen3.5-4B
fine-tuned on 200K grounding examples from
DeskForge-1M.
Given a desktop screenshot and a natural-language instruction, it answers with
the next action as a pyautogui call whose coordinates are fractions of the
screen.
Project page · Dataset · Code · Paper
Grounding on DeskForge-1M held-out conditions (accuracy, %). New Scenes are unseen scenes with seen attributes; Theme, App and Resolution hold an appearance preset, three applications and a display resolution out of training entirely.
| Model | New Scenes | Theme | App | Resolution | Mean |
|---|---|---|---|---|---|
| UI-TARS-1.5-7B | 67.68 | 64.96 | 66.68 | 61.67 | 65.25 |
| GroundNext-7B | 73.57 | 67.67 | 72.10 | 68.87 | 70.55 |
| UI-Venus-2-9B | 81.45 | 77.44 | 79.40 | 77.40 | 78.92 |
| Qwen3.5-4B | 78.56 | 74.03 | 76.11 | 76.34 | 76.26 |
| DeskForge-Qwen3.5-4B | 90.14 | 87.36 | 87.50 | 85.15 | 87.54 |
Grounding on external GUI benchmarks (accuracy, %), none of which is used for training:
| Model | ScreenSpot-Pro | ScreenSpot-v2 | OSWorld-G | UI-Vision | MMBench-GUI |
|---|---|---|---|---|---|
| Qwen3.5-4B | 30.17 | 80.11 | 51.24 | 17.24 | 61.83 |
| DeskForge-Qwen3.5-4B | 41.68 | 90.33 | 61.35 | 28.77 | 76.82 |
Long-horizon task completion (tasks solved). A fixed Qwen3.6-27B planner decides each step and the action model locates its targets, so only the action model differs between the rows:
| Action model | WebArena-Infinity (119 tasks) | OpenApps (100 tasks) |
|---|---|---|
| Qwen3.5-4B | 31 | 3 |
| DeskForge-Qwen3.5-4B | 50 | 15 |
The model takes the system prompt it was trained with, one screenshot, and the task.
import re
from PIL import Image
MODEL = "docling-project/DeskForge-Qwen3.5-4B"
SYSTEM_PROMPT = """You are a computer-use agent operating a desktop graphical interface. At each step you see the user's task, a screenshot of the current screen, and the actions you have already taken. Reply with the next action as pyautogui code and nothing else -- no explanation, no code fence, no commentary.
Coordinates are fractions of the screen, not pixels: x runs from 0.0 at the left edge to 1.0 at the right edge, y from 0.0 at the top to 1.0 at the bottom. Write both with four decimals.
These are the only actions available:
pyautogui.click(x=0.0000, y=0.0000)
pyautogui.doubleClick(x=0.0000, y=0.0000)
pyautogui.rightClick(x=0.0000, y=0.0000)
pyautogui.middleClick(x=0.0000, y=0.0000)
computer.tripleClick(x=0.0000, y=0.0000)
pyautogui.moveTo(x=0.0000, y=0.0000)
pyautogui.dragTo(x=0.0000, y=0.0000, button='left')
pyautogui.scroll(-4)
pyautogui.hscroll(4)
pyautogui.write(message='text to type')
pyautogui.press('enter')
pyautogui.hotkey(['ctrl', 'c'])
computer.wait()
computer.terminate(status='success')
To scroll at a particular place, move there first and then scroll. When the task is finished, or cannot be finished, end with computer.terminate."""
def fit(image, max_pixels=2_097_152):
"""Downscale to at most max_pixels (even sides, Lanczos), as in training."""
w, h = image.size
if w * h <= max_pixels:
return image
s = (max_pixels / (w * h)) ** 0.5
return image.resize((max(2, int(w * s) // 2 * 2), max(2, int(h * s) // 2 * 2)), Image.LANCZOS)
def user_text(instruction):
return f"Task: {instruction}\n\nActions already taken:\n(none -- this is the first step)\n\nNext action:"
def chat(instruction):
return [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": user_text(instruction)}]},
]
def to_pixels(action, screenshot):
x, y = map(float, re.search(r"x=([\d.]+), y=([\d.]+)", action).groups())
return round(x * screenshot.width), round(y * screenshot.height)
screenshot = Image.open("screenshot.png").convert("RGB")
instruction = "Open the File menu of the text editor."
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(MODEL, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(MODEL)
prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
add_generation_prompt=True, enable_thinking=False)
inputs = processor(text=[prompt], images=[fit(screenshot)], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
action = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()
print(action, to_pixels(action, screenshot)) # e.g. pyautogui.click(x=0.0412, y=0.0535) (105, 77)
from transformers import AutoProcessor
from vllm import LLM, SamplingParams
processor = AutoProcessor.from_pretrained(MODEL)
llm = LLM(model=MODEL, max_model_len=16384, limit_mm_per_prompt={"image": 1})
prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
add_generation_prompt=True, enable_thinking=False)
outputs = llm.generate({"prompt": prompt, "multi_modal_data": {"image": fit(screenshot)}},
SamplingParams(temperature=0.0, max_tokens=128))
action = outputs[0].outputs[0].text.strip()
print(action, to_pixels(action, screenshot))
vllm serve docling-project/DeskForge-Qwen3.5-4B --max-model-len 16384
import base64
import io
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
buffer = io.BytesIO()
fit(screenshot).save(buffer, format="PNG")
image_url = "data:image/png;base64," + base64.b64encode(buffer.getvalue()).decode()
response = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": image_url}},
{"type": "text", "text": user_text(instruction)},
]},
],
max_tokens=128,
temperature=0.0,
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
action = response.choices[0].message.content.strip()
print(action, to_pixels(action, screenshot))
| Base model | Qwen/Qwen3.5-4B |
| Data | 200,000 single-target grounding examples from the DeskForge-1M training split: a screenshot, an instruction, and a click on the target |
| Method | LoRA on the language model (rank 8, α 32, dropout 0.05), vision encoder and aligner frozen, merged into the base weights |
| Optimization | One epoch, 3,125 updates, effective batch size 64, AdamW, learning rate 10⁻⁴ with cosine decay and 3% warm-up, BF16 |
| Inputs | Screenshots capped at 2,097,152 pixels; sequences up to 4,096 tokens |
| Hardware | 8 × NVIDIA H100 80GB |
The model predicts the next GUI action, mainly where to click, from a screenshot and an instruction, and serves as the action model under a planner in multi-step tasks.
Apache 2.0, following the base model; see LICENSE.
@misc{gurbuz2026deskforge,
title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
year={2026},
eprint={2610.02320},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02320},
}