deskmind-ai/bench

Real-desktop benchmark for computer-use agents: sandbox macOS tasks, graders, reference runs · 得心

Python

0

5 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

What went wrong with confidence routing in my local 0.8B/4B agent (r/LocalLLM)

I've been building a local computer-use agent for macOS over the last three weeks. It uses Qwen3.5 0.8B and 4B, fine-tuned with LoRA and running in 8-bit MLX. The training runs and the mistakes below are mine. I ran into a problem with label smoothing that I thought might be worth sharing. The…

2

Oct 6, 2026

README

DeskMind 得心 — 得心,应手。

DeskMind
DeskMind Bench · Sandbox macOS desktop tasks, their graders, and reference runs.
中文 · Tasks · Scoring · Results · Versions

Part of DeskMind: start there for the app, the demo and the other components.

Issues and questions go to deskmind-ai/deskmind/issues. Pull requests come here.


DeskMind · 得心 — 得心,应手。 (from 得心应手: what the mind decides, the hand carries out) is a family of open-source projects that let an agent see your screen, decide the next step, and act on the real desktop, all locally.

This repository is the Bench: the task suite our desktop numbers are measured on, the graders that score it, and the scorer of record. Every DeskMind result names a suite hash and a harness version from this repository.

RepositoryRole
👁deskmind-ai/eyesfind the target on a screenshot
🧠deskmind-ai/braindecide the next step, with calibrated confidence
✋deskmind-ai/handsdrive the real macOS desktop
📐deskmind-ai/benchsandbox desktop tasks and graders to reproduce our numbers

What the diag suite tests

Thirteen tasks on the real Finder, TextEdit and Safari, each in a sandbox folder unpacked from a fixture. Each one isolates one thing a desktop agent gets wrong (docs/tasks.md):

  • Finder operations: make a folder, move a file, navigate down two levels and back up, sort files into a new folder, rename several files in turn.
  • Editing text exactly: change two fields and save; type Chinese with full-width punctuation character for character; find one line in a 60-line log that needs scrolling.
  • Crossing apps: read a table in Safari and append it sorted to a CSV; copy values from one document to another.
  • Behaving well: ask when the goal is ambiguous instead of guessing; edit only the named one of two near-identical documents; stop when cancelled halfway.

A run passes (strict) only if every checkpoint passes, no guard is broken, no forbidden side effect happened, and the sentinel files are untouched. Partial credit is reported separately and never replaces strict.

Run and score

Scoring needs only Python and PyYAML: no Mac, no driver, no model.

git clone https://github.com/deskmind-ai/bench && cd bench
pip install -e .

deskmind-bench verify                  # every task: its effect oracle passes, doing nothing fails
deskmind-bench hash                    # 5eec62a0c662 -- the suite hash of the reference results
deskmind-bench versions                # harness versions and their commits
deskmind-bench score path/to/runs/ --suite-version diag-v25 --out mine.json   # regrade finished runs from disk
deskmind-bench table mine.json         # per-task passes as a markdown table

score reads each run's workspace, record and trace as DeskMind Hands writes them, grades them again with this repository's graders, and reports strict and partial pass, false DONE, DONE timing and step latency. It never imports the driver.

Executing runs needs a Mac with DeskMind Hands, Peekaboo and its permissions, and a planner behind /v1/systemone:

git clone https://github.com/deskmind-ai/hands ../hands   # next to this checkout, as hands expects bench
pip install -e ../hands -e ".[run]"    # adds deskmind-hands
deskmind-bench run --set diag --repeats 3 --url http://127.0.0.1:8793 --label my-planner --out my-planner.json

The runner refuses a non-local URL unless you pass --allow-remote, and removes a hosted API key from the environment unless you ask for the hosted reference (--hosted).

Reference results

Real macOS desktop, 13 tasks × 3 runs, projection layer on, strict pass. Full per-task tables: results/reference.md (and reference.json).

harnessconfigstrict passfalse DONEstep p50
v25DeskMind Brain router (0.8B → 4B), g18b, 8-bit, threshold 0.96 (current)39/39 (100%)02.85 s¹
v25DeskMind Brain router (0.8B → 4B), 4B g14, 8-bit36/39 (92%)00.57 s
v25DeskMind Brain router, 4B g17, 8-bit36/39 (92%)3n/a
v23DeskMind Brain router, 4B g14, 8-bit35/38 (92%), 1 env error00.59 s
v23Jev (TypeSafe AI, cloud reference)33/38 (87%), 1 env error20.36 s
v21DeskMind Brain router (0.8B → 4B), rev. v7b32/39 (82%)23.3 s
v20Jev (TypeSafe AI, cloud reference)33/39 (85%)00.4 s
v20DeskMind Brain router, rev. v729/39 (74%)33.2 s
v19DeskMind Brain router, rev. v730/39 (77%)23.6 s
v19DeskMind Brain 4B, g11b27/39 (69%)33.4 s
v19DeskMind Brain 4B, g10b24/39 (62%)7–
v19Jev (TypeSafe AI, cloud reference)31/37 (84%), 2 env errors21.0 s

¹ Run through the DeskMind app, whose step time includes the app's own checks and the steps handed to the second model; the other rows ran from the command line. p95 9.82 s.

Versions

The tasks, fixtures and graders are the same in all versions (suite hash 5eec62a0c662). What changed is the harness: how the desktop is shown to the planner and how actions are carried out. results/versions.md (and versions.json) lists what changed in each.

The hands commits below are from DeskMind Hands' development history, which was rebuilt into a single public commit on 2026-10-02: they name each version but cannot be checked out, so the older reference numbers cannot be rerun exactly. The public DeskMind Hands contains every change listed. New reference runs are made on it and published under a new harness version, with its commit.

suite versionhands commitsuite hashstatus
diag-v193aee9845eec62a0c662frozen
diag-v207934cfa5eec62a0c662superseded
diag-v21f1df1185eec62a0c662frozen
diag-v221c3f47d5eec62a0c662frozen
diag-v23713961b5eec62a0c662frozen
diag-v244d120335eec62a0c662frozen
diag-v253d544925eec62a0c662current

Read the numbers with care

  • n = 3 per task. Thirteen tasks × three runs is 39 runs. One run is 2.6 points, and a difference of one or two runs between configs is noise. Thirteen tasks is the thinnest part of this benchmark: help grow it with Add a task in 30 minutes.
  • Runs cluster by task. Almost every task passes 3/3 or 0/3, so the effective sample is closer to 13 tasks than to 39 runs. Read the per-task table, not only the total, and use a task-level (clustered) bootstrap for intervals.
  • Results depend on the harness version. The same planner scored 30/39 on v19 and 29/39 on v20 with identical tasks. Compare only within one version.
  • macOS only, on one machine (Apple M4 Pro, macOS 27, zh-Hans locale, Peekaboo 4.3.0). The tasks are written in Chinese; other locales and OS versions are untested.
  • G03 (read a table from a web page) and G05 (ask before acting) were unsolved by every config up to v21. The v25 g14 router passed every task except G03; the v25 g18b router passes all thirteen, 3/3 each. With n = 3 that is not proof that G03 is solved for good: treat it as "no longer failing on this suite".
  • The suite is public. Anyone can train on these tasks. Our own training tasks are generated separately and never include them.

License

Apache-2.0, see LICENSE and NOTICE. The DeskMind name, 得心, the logo and Xiaofang are not covered by the code licence. You may use them to refer to the project, but not in modified form or to imply endorsement.

agents
benchmark
computer-use
desktop-automation
evaluation
macos

deskmind-ai/bench

Real-desktop benchmark for computer-use agents: sandbox macOS tasks, graders, reference runs · 得心

Python

0

5 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

What went wrong with confidence routing in my local 0.8B/4B agent (r/LocalLLM)

I've been building a local computer-use agent for macOS over the last three weeks. It uses Qwen3.5 0.8B and 4B, fine-tuned with LoRA and running in 8-bit MLX. The training runs and the mistakes below are mine. I ran into a problem with label smoothing that I thought might be worth sharing. The…

2

Oct 6, 2026

README

DeskMind 得心 — 得心,应手。

DeskMind
DeskMind Bench · Sandbox macOS desktop tasks, their graders, and reference runs.
中文 · Tasks · Scoring · Results · Versions

Part of DeskMind: start there for the app, the demo and the other components.

Issues and questions go to deskmind-ai/deskmind/issues. Pull requests come here.


DeskMind · 得心 — 得心,应手。 (from 得心应手: what the mind decides, the hand carries out) is a family of open-source projects that let an agent see your screen, decide the next step, and act on the real desktop, all locally.

This repository is the Bench: the task suite our desktop numbers are measured on, the graders that score it, and the scorer of record. Every DeskMind result names a suite hash and a harness version from this repository.

RepositoryRole
👁deskmind-ai/eyesfind the target on a screenshot
🧠deskmind-ai/braindecide the next step, with calibrated confidence
✋deskmind-ai/handsdrive the real macOS desktop
📐deskmind-ai/benchsandbox desktop tasks and graders to reproduce our numbers

What the diag suite tests

Thirteen tasks on the real Finder, TextEdit and Safari, each in a sandbox folder unpacked from a fixture. Each one isolates one thing a desktop agent gets wrong (docs/tasks.md):

  • Finder operations: make a folder, move a file, navigate down two levels and back up, sort files into a new folder, rename several files in turn.
  • Editing text exactly: change two fields and save; type Chinese with full-width punctuation character for character; find one line in a 60-line log that needs scrolling.
  • Crossing apps: read a table in Safari and append it sorted to a CSV; copy values from one document to another.
  • Behaving well: ask when the goal is ambiguous instead of guessing; edit only the named one of two near-identical documents; stop when cancelled halfway.

A run passes (strict) only if every checkpoint passes, no guard is broken, no forbidden side effect happened, and the sentinel files are untouched. Partial credit is reported separately and never replaces strict.

Run and score

Scoring needs only Python and PyYAML: no Mac, no driver, no model.

git clone https://github.com/deskmind-ai/bench && cd bench
pip install -e .

deskmind-bench verify                  # every task: its effect oracle passes, doing nothing fails
deskmind-bench hash                    # 5eec62a0c662 -- the suite hash of the reference results
deskmind-bench versions                # harness versions and their commits
deskmind-bench score path/to/runs/ --suite-version diag-v25 --out mine.json   # regrade finished runs from disk
deskmind-bench table mine.json         # per-task passes as a markdown table

score reads each run's workspace, record and trace as DeskMind Hands writes them, grades them again with this repository's graders, and reports strict and partial pass, false DONE, DONE timing and step latency. It never imports the driver.

Executing runs needs a Mac with DeskMind Hands, Peekaboo and its permissions, and a planner behind /v1/systemone:

git clone https://github.com/deskmind-ai/hands ../hands   # next to this checkout, as hands expects bench
pip install -e ../hands -e ".[run]"    # adds deskmind-hands
deskmind-bench run --set diag --repeats 3 --url http://127.0.0.1:8793 --label my-planner --out my-planner.json

The runner refuses a non-local URL unless you pass --allow-remote, and removes a hosted API key from the environment unless you ask for the hosted reference (--hosted).

Reference results

Real macOS desktop, 13 tasks × 3 runs, projection layer on, strict pass. Full per-task tables: results/reference.md (and reference.json).

harnessconfigstrict passfalse DONEstep p50
v25DeskMind Brain router (0.8B → 4B), g18b, 8-bit, threshold 0.96 (current)39/39 (100%)02.85 s¹
v25DeskMind Brain router (0.8B → 4B), 4B g14, 8-bit36/39 (92%)00.57 s
v25DeskMind Brain router, 4B g17, 8-bit36/39 (92%)3n/a
v23DeskMind Brain router, 4B g14, 8-bit35/38 (92%), 1 env error00.59 s
v23Jev (TypeSafe AI, cloud reference)33/38 (87%), 1 env error20.36 s
v21DeskMind Brain router (0.8B → 4B), rev. v7b32/39 (82%)23.3 s
v20Jev (TypeSafe AI, cloud reference)33/39 (85%)00.4 s
v20DeskMind Brain router, rev. v729/39 (74%)33.2 s
v19DeskMind Brain router, rev. v730/39 (77%)23.6 s
v19DeskMind Brain 4B, g11b27/39 (69%)33.4 s
v19DeskMind Brain 4B, g10b24/39 (62%)7–
v19Jev (TypeSafe AI, cloud reference)31/37 (84%), 2 env errors21.0 s

¹ Run through the DeskMind app, whose step time includes the app's own checks and the steps handed to the second model; the other rows ran from the command line. p95 9.82 s.

Versions

The tasks, fixtures and graders are the same in all versions (suite hash 5eec62a0c662). What changed is the harness: how the desktop is shown to the planner and how actions are carried out. results/versions.md (and versions.json) lists what changed in each.

The hands commits below are from DeskMind Hands' development history, which was rebuilt into a single public commit on 2026-10-02: they name each version but cannot be checked out, so the older reference numbers cannot be rerun exactly. The public DeskMind Hands contains every change listed. New reference runs are made on it and published under a new harness version, with its commit.

suite versionhands commitsuite hashstatus
diag-v193aee9845eec62a0c662frozen
diag-v207934cfa5eec62a0c662superseded
diag-v21f1df1185eec62a0c662frozen
diag-v221c3f47d5eec62a0c662frozen
diag-v23713961b5eec62a0c662frozen
diag-v244d120335eec62a0c662frozen
diag-v253d544925eec62a0c662current

Read the numbers with care

  • n = 3 per task. Thirteen tasks × three runs is 39 runs. One run is 2.6 points, and a difference of one or two runs between configs is noise. Thirteen tasks is the thinnest part of this benchmark: help grow it with Add a task in 30 minutes.
  • Runs cluster by task. Almost every task passes 3/3 or 0/3, so the effective sample is closer to 13 tasks than to 39 runs. Read the per-task table, not only the total, and use a task-level (clustered) bootstrap for intervals.
  • Results depend on the harness version. The same planner scored 30/39 on v19 and 29/39 on v20 with identical tasks. Compare only within one version.
  • macOS only, on one machine (Apple M4 Pro, macOS 27, zh-Hans locale, Peekaboo 4.3.0). The tasks are written in Chinese; other locales and OS versions are untested.
  • G03 (read a table from a web page) and G05 (ask before acting) were unsolved by every config up to v21. The v25 g14 router passed every task except G03; the v25 g18b router passes all thirteen, 3/3 each. With n = 3 that is not proof that G03 is solved for good: treat it as "no longer failing on this suite".
  • The suite is public. Anyone can train on these tasks. Our own training tasks are generated separately and never include them.

License

Apache-2.0, see LICENSE and NOTICE. The DeskMind name, 得心, the logo and Xiaofang are not covered by the code licence. You may use them to refer to the project, but not in modified form or to imply endorsement.

agents
benchmark
computer-use
desktop-automation
evaluation
macos