docling-project/DeskForge-Qwen3.5-4B

Model

DeskForge-Qwen3.5-4B

0

5 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

DeskForge-Qwen3.5-4B

DeskForge-Qwen3.5-4B is Qwen3.5-4B fine-tuned on 200K grounding examples from DeskForge-1M. Given a desktop screenshot and a natural-language instruction, it answers with the next action as a pyautogui call whose coordinates are fractions of the screen.

Project page · Dataset · Code · Paper

Results

Grounding on DeskForge-1M held-out conditions (accuracy, %). New Scenes are unseen scenes with seen attributes; Theme, App and Resolution hold an appearance preset, three applications and a display resolution out of training entirely.

ModelNew ScenesThemeAppResolutionMean
UI-TARS-1.5-7B67.6864.9666.6861.6765.25
GroundNext-7B73.5767.6772.1068.8770.55
UI-Venus-2-9B81.4577.4479.4077.4078.92
Qwen3.5-4B78.5674.0376.1176.3476.26
DeskForge-Qwen3.5-4B90.1487.3687.5085.1587.54

Grounding on external GUI benchmarks (accuracy, %), none of which is used for training:

ModelScreenSpot-ProScreenSpot-v2OSWorld-GUI-VisionMMBench-GUI
Qwen3.5-4B30.1780.1151.2417.2461.83
DeskForge-Qwen3.5-4B41.6890.3361.3528.7776.82

Long-horizon task completion (tasks solved). A fixed Qwen3.6-27B planner decides each step and the action model locates its targets, so only the action model differs between the rows:

Action modelWebArena-Infinity (119 tasks)OpenApps (100 tasks)
Qwen3.5-4B313
DeskForge-Qwen3.5-4B5015

Usage

The model takes the system prompt it was trained with, one screenshot, and the task.

import re

from PIL import Image

MODEL = "docling-project/DeskForge-Qwen3.5-4B"

SYSTEM_PROMPT = """You are a computer-use agent operating a desktop graphical interface. At each step you see the user's task, a screenshot of the current screen, and the actions you have already taken. Reply with the next action as pyautogui code and nothing else -- no explanation, no code fence, no commentary.

Coordinates are fractions of the screen, not pixels: x runs from 0.0 at the left edge to 1.0 at the right edge, y from 0.0 at the top to 1.0 at the bottom. Write both with four decimals.

These are the only actions available:
pyautogui.click(x=0.0000, y=0.0000)
pyautogui.doubleClick(x=0.0000, y=0.0000)
pyautogui.rightClick(x=0.0000, y=0.0000)
pyautogui.middleClick(x=0.0000, y=0.0000)
computer.tripleClick(x=0.0000, y=0.0000)
pyautogui.moveTo(x=0.0000, y=0.0000)
pyautogui.dragTo(x=0.0000, y=0.0000, button='left')
pyautogui.scroll(-4)
pyautogui.hscroll(4)
pyautogui.write(message='text to type')
pyautogui.press('enter')
pyautogui.hotkey(['ctrl', 'c'])
computer.wait()
computer.terminate(status='success')

To scroll at a particular place, move there first and then scroll. When the task is finished, or cannot be finished, end with computer.terminate."""


def fit(image, max_pixels=2_097_152):
    """Downscale to at most max_pixels (even sides, Lanczos), as in training."""
    w, h = image.size
    if w * h <= max_pixels:
        return image
    s = (max_pixels / (w * h)) ** 0.5
    return image.resize((max(2, int(w * s) // 2 * 2), max(2, int(h * s) // 2 * 2)), Image.LANCZOS)


def user_text(instruction):
    return f"Task: {instruction}\n\nActions already taken:\n(none -- this is the first step)\n\nNext action:"


def chat(instruction):
    return [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": user_text(instruction)}]},
    ]


def to_pixels(action, screenshot):
    x, y = map(float, re.search(r"x=([\d.]+), y=([\d.]+)", action).groups())
    return round(x * screenshot.width), round(y * screenshot.height)


screenshot = Image.open("screenshot.png").convert("RGB")
instruction = "Open the File menu of the text editor."
Transformers
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(MODEL, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(MODEL)

prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
                                       add_generation_prompt=True, enable_thinking=False)
inputs = processor(text=[prompt], images=[fit(screenshot)], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
action = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()

print(action, to_pixels(action, screenshot))  # e.g. pyautogui.click(x=0.0412, y=0.0535) (105, 77)
vLLM
from transformers import AutoProcessor
from vllm import LLM, SamplingParams

processor = AutoProcessor.from_pretrained(MODEL)
llm = LLM(model=MODEL, max_model_len=16384, limit_mm_per_prompt={"image": 1})

prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
                                       add_generation_prompt=True, enable_thinking=False)
outputs = llm.generate({"prompt": prompt, "multi_modal_data": {"image": fit(screenshot)}},
                       SamplingParams(temperature=0.0, max_tokens=128))
action = outputs[0].outputs[0].text.strip()

print(action, to_pixels(action, screenshot))
vLLM server (OpenAI-compatible API)
vllm serve docling-project/DeskForge-Qwen3.5-4B --max-model-len 16384
import base64
import io

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

buffer = io.BytesIO()
fit(screenshot).save(buffer, format="PNG")
image_url = "data:image/png;base64," + base64.b64encode(buffer.getvalue()).decode()

response = client.chat.completions.create(
    model=MODEL,
    messages=[
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": [
            {"type": "image_url", "image_url": {"url": image_url}},
            {"type": "text", "text": user_text(instruction)},
        ]},
    ],
    max_tokens=128,
    temperature=0.0,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
action = response.choices[0].message.content.strip()

print(action, to_pixels(action, screenshot))

Training

Base modelQwen/Qwen3.5-4B
Data200,000 single-target grounding examples from the DeskForge-1M training split: a screenshot, an instruction, and a click on the target
MethodLoRA on the language model (rank 8, α 32, dropout 0.05), vision encoder and aligner frozen, merged into the base weights
OptimizationOne epoch, 3,125 updates, effective batch size 64, AdamW, learning rate 10⁻⁴ with cosine decay and 3% warm-up, BF16
InputsScreenshots capped at 2,097,152 pixels; sequences up to 4,096 tokens
Hardware8 × NVIDIA H100 80GB

Intended use and limitations

The model predicts the next GUI action, mainly where to click, from a screenshot and an instruction, and serves as the action model under a planner in multi-step tasks.

  • Training screenshots come from Linux (Xfce) desktops with several applications per screen, overlapping windows, and appearance presets that include Windows- and macOS-inspired styles.
  • Training targets are single clicks; other action types and task planning come from the surrounding agent.
  • The model is not a safe autonomous agent. Review its actions before running them on real systems.

License

Apache 2.0, following the base model; see LICENSE.

Citation

@misc{gurbuz2026deskforge,
      title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
      author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
      year={2026},
      eprint={2610.02320},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02320},
}
computer-use
conversational
desktop
endpoints_compatible
gui-agent
gui-grounding
image-text-to-text
qwen3_5
safetensors
screen-understanding
transformers

docling-project/DeskForge-Qwen3.5-4B

Model

DeskForge-Qwen3.5-4B

0

5 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

DeskForge-Qwen3.5-4B

DeskForge-Qwen3.5-4B is Qwen3.5-4B fine-tuned on 200K grounding examples from DeskForge-1M. Given a desktop screenshot and a natural-language instruction, it answers with the next action as a pyautogui call whose coordinates are fractions of the screen.

Project page · Dataset · Code · Paper

Results

Grounding on DeskForge-1M held-out conditions (accuracy, %). New Scenes are unseen scenes with seen attributes; Theme, App and Resolution hold an appearance preset, three applications and a display resolution out of training entirely.

ModelNew ScenesThemeAppResolutionMean
UI-TARS-1.5-7B67.6864.9666.6861.6765.25
GroundNext-7B73.5767.6772.1068.8770.55
UI-Venus-2-9B81.4577.4479.4077.4078.92
Qwen3.5-4B78.5674.0376.1176.3476.26
DeskForge-Qwen3.5-4B90.1487.3687.5085.1587.54

Grounding on external GUI benchmarks (accuracy, %), none of which is used for training:

ModelScreenSpot-ProScreenSpot-v2OSWorld-GUI-VisionMMBench-GUI
Qwen3.5-4B30.1780.1151.2417.2461.83
DeskForge-Qwen3.5-4B41.6890.3361.3528.7776.82

Long-horizon task completion (tasks solved). A fixed Qwen3.6-27B planner decides each step and the action model locates its targets, so only the action model differs between the rows:

Action modelWebArena-Infinity (119 tasks)OpenApps (100 tasks)
Qwen3.5-4B313
DeskForge-Qwen3.5-4B5015

Usage

The model takes the system prompt it was trained with, one screenshot, and the task.

import re

from PIL import Image

MODEL = "docling-project/DeskForge-Qwen3.5-4B"

SYSTEM_PROMPT = """You are a computer-use agent operating a desktop graphical interface. At each step you see the user's task, a screenshot of the current screen, and the actions you have already taken. Reply with the next action as pyautogui code and nothing else -- no explanation, no code fence, no commentary.

Coordinates are fractions of the screen, not pixels: x runs from 0.0 at the left edge to 1.0 at the right edge, y from 0.0 at the top to 1.0 at the bottom. Write both with four decimals.

These are the only actions available:
pyautogui.click(x=0.0000, y=0.0000)
pyautogui.doubleClick(x=0.0000, y=0.0000)
pyautogui.rightClick(x=0.0000, y=0.0000)
pyautogui.middleClick(x=0.0000, y=0.0000)
computer.tripleClick(x=0.0000, y=0.0000)
pyautogui.moveTo(x=0.0000, y=0.0000)
pyautogui.dragTo(x=0.0000, y=0.0000, button='left')
pyautogui.scroll(-4)
pyautogui.hscroll(4)
pyautogui.write(message='text to type')
pyautogui.press('enter')
pyautogui.hotkey(['ctrl', 'c'])
computer.wait()
computer.terminate(status='success')

To scroll at a particular place, move there first and then scroll. When the task is finished, or cannot be finished, end with computer.terminate."""


def fit(image, max_pixels=2_097_152):
    """Downscale to at most max_pixels (even sides, Lanczos), as in training."""
    w, h = image.size
    if w * h <= max_pixels:
        return image
    s = (max_pixels / (w * h)) ** 0.5
    return image.resize((max(2, int(w * s) // 2 * 2), max(2, int(h * s) // 2 * 2)), Image.LANCZOS)


def user_text(instruction):
    return f"Task: {instruction}\n\nActions already taken:\n(none -- this is the first step)\n\nNext action:"


def chat(instruction):
    return [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": [{"type": "image"}, {"type": "text", "text": user_text(instruction)}]},
    ]


def to_pixels(action, screenshot):
    x, y = map(float, re.search(r"x=([\d.]+), y=([\d.]+)", action).groups())
    return round(x * screenshot.width), round(y * screenshot.height)


screenshot = Image.open("screenshot.png").convert("RGB")
instruction = "Open the File menu of the text editor."
Transformers
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(MODEL, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(MODEL)

prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
                                       add_generation_prompt=True, enable_thinking=False)
inputs = processor(text=[prompt], images=[fit(screenshot)], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
action = processor.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()

print(action, to_pixels(action, screenshot))  # e.g. pyautogui.click(x=0.0412, y=0.0535) (105, 77)
vLLM
from transformers import AutoProcessor
from vllm import LLM, SamplingParams

processor = AutoProcessor.from_pretrained(MODEL)
llm = LLM(model=MODEL, max_model_len=16384, limit_mm_per_prompt={"image": 1})

prompt = processor.apply_chat_template(chat(instruction), tokenize=False,
                                       add_generation_prompt=True, enable_thinking=False)
outputs = llm.generate({"prompt": prompt, "multi_modal_data": {"image": fit(screenshot)}},
                       SamplingParams(temperature=0.0, max_tokens=128))
action = outputs[0].outputs[0].text.strip()

print(action, to_pixels(action, screenshot))
vLLM server (OpenAI-compatible API)
vllm serve docling-project/DeskForge-Qwen3.5-4B --max-model-len 16384
import base64
import io

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

buffer = io.BytesIO()
fit(screenshot).save(buffer, format="PNG")
image_url = "data:image/png;base64," + base64.b64encode(buffer.getvalue()).decode()

response = client.chat.completions.create(
    model=MODEL,
    messages=[
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": [
            {"type": "image_url", "image_url": {"url": image_url}},
            {"type": "text", "text": user_text(instruction)},
        ]},
    ],
    max_tokens=128,
    temperature=0.0,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
action = response.choices[0].message.content.strip()

print(action, to_pixels(action, screenshot))

Training

Base modelQwen/Qwen3.5-4B
Data200,000 single-target grounding examples from the DeskForge-1M training split: a screenshot, an instruction, and a click on the target
MethodLoRA on the language model (rank 8, α 32, dropout 0.05), vision encoder and aligner frozen, merged into the base weights
OptimizationOne epoch, 3,125 updates, effective batch size 64, AdamW, learning rate 10⁻⁴ with cosine decay and 3% warm-up, BF16
InputsScreenshots capped at 2,097,152 pixels; sequences up to 4,096 tokens
Hardware8 × NVIDIA H100 80GB

Intended use and limitations

The model predicts the next GUI action, mainly where to click, from a screenshot and an instruction, and serves as the action model under a planner in multi-step tasks.

  • Training screenshots come from Linux (Xfce) desktops with several applications per screen, overlapping windows, and appearance presets that include Windows- and macOS-inspired styles.
  • Training targets are single clicks; other action types and task planning come from the surrounding agent.
  • The model is not a safe autonomous agent. Review its actions before running them on real systems.

License

Apache 2.0, following the base model; see LICENSE.

Citation

@misc{gurbuz2026deskforge,
      title={DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents},
      author={A. Said Gurbuz and Ahmed Nassar and Sunghwan Hong and Marc Pollefeys and Peter W. J. Staar},
      year={2026},
      eprint={2610.02320},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02320},
}
computer-use
conversational
desktop
endpoints_compatible
gui-agent
gui-grounding
image-text-to-text
qwen3_5
safetensors
screen-understanding
transformers