Small enough to run on your Mac. Smart enough to ask. An open-source computer use agent for macOS, with local models (MLX) that see, decide and act. 得心,应手。
See the code
Small enough to run on your Mac. Smart enough to ask.
Open-source computer use for your Mac. DeskMind reads the screen, works out the next step and acts, all with small models running on your Mac. When a task could mean two things, it asks you instead of guessing.
中文 · Website · Docs · Download for Mac · Models · Discussions · Roadmap · Changelog
Watch the 56-second demo on deskmind.dev: a real recording with the released model. Two orders match "Lisa Wong", so it asks which one before writing.
How it was built, and what went wrong on the way: 3 weeks, 20 training rounds, $600. If DeskMind is useful or interesting to you, a star on this repo helps other people find it.
A small model, on your Mac. A 0.8B model decides each step and hands the unsure ones to a 4B. Both run on your Mac: no cloud round-trip, no per-step bill. The 0.8B decides in about 0.5 s; the 4B answers in about 3.6 s (median decision times).
System One: choices, not guesses. Each step is a multiple-choice question. The model scores every option instead of writing text, so every option gets a probability. Unsure steps go to the 4B or to you, and any agent can call it through POST /v1/systemone. In 39 real-desktop runs it never said "done" when the task was not done.
Open, from eyes to hands. Eyes, Brain, Hands and the Mac app are open source, along with Bench, which grades them. Every result below comes with its sample size, so you can reproduce it.
Use the Mac app. Download DeskMind for Mac (macOS 15+, Apple Silicon, signed and notarized). It ships the G18b release models. On first run it downloads the models (about 5.3 GB) and walks you through the permissions: Install the app.
Or run the model yourself. This runs the 4B alone (release G18b). It answers one step of a desktop task; it does not drive the desktop by itself.
# Terminal 1: get Brain and the model, then serve it
git clone https://github.com/deskmind-ai/brain && cd brain && uv sync --extra mlx
uv run hf download deskmind/brain-4b --revision g18b-q8 --local-dir models/brain-4b
uv run deskmind-brain-serve --predictor mlx:models/brain-4b --port 8793 --two-stage
The 4B is about 4.5 GB to download. If the download fails with a CAS Client Error, rerun it with HF_HUB_DISABLE_XET=1 in front. In mainland China, ModelScope carries the same files:
uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b.
# Terminal 2 (in the same brain folder): ask for the next step of a real desktop task
curl -s localhost:8793/v1/systemone -H 'Content-Type: application/json' -d @examples/request.json
Success looks like a typed decision with a probability for every option, e.g.
{"answers": {"operation": {"choice": "CLICK", "probabilities": {"CLICK": 0.96, "OPEN": 0.005, …}}, "click_target": {…}}}.
A probability is the model's weighting of the options, not a guarantee that the step is right.
The router (0.8B → 4B), as released: also download deskmind/brain-0.8b --revision g18b-q8 into models/brain-0.8b (0.8 GB), then serve
uv run deskmind-brain-serve --predictor mlx:models/brain-0.8b --escalate-to mlx:models/brain-4b --two-stage --port 8796.
The threshold (0.96) ships with the weights; each reply adds a routing record such as {"by": "strong", "reason": "low_conf", "fast_conf": 0.956}.
Step by step, with the full reply explained: Quickstart. To drive the desktop, add Hands.
| Role | ||
|---|---|---|
| Eyes | finds the target on screen | 4B visual grounder, for apps without an accessibility tree |
| Brain | decides the next step | 0.8B and 4B, MLX, 0.8B → 4B routing, /v1/systemone |
| Hands | observes and acts on macOS | accessibility and vision modes, budgets, cancellation |
| App (this repository) | brings it to your Mac | native app with a background helper; download |
| Bench | checks what really happened | sandbox desktop tasks with strict final-state graders |
Models: huggingface.co/deskmind (brain-0.8b, brain-4b, eyes-4b). Website: deskmind.dev. Docs: deskmind.dev/docs.
| Path | What |
|---|---|
app/ | The Mac app's source (Swift) and its build; build it yourself |
docs/ | How the app, its helper and the local model servers fit together |
brand/ | Logo, Xiaofang and social images; rules in BRAND.md |
.github/workflows/app.yml | CI: every app change is built and tested; a version tag is signed, notarized and drafted as a release |
ROADMAP.md | What we are working on next |
Mac app releases are published here: Releases; what changed in each one, in short: CHANGELOG.md.
| What | Setting | Result |
|---|---|---|
| Real-desktop tasks | Bench v25, 13 tasks × 3 runs, strict graders; router G18b (0.8B → 4B, 8-bit, threshold 0.96), through the app, one M4 Pro (48 GB) | 39/39 passed; 0 false "done" |
| Decision time | the same 39 runs, 208 decisions | median 0.48 s when the 0.8B answers (about 30% of steps), 3.6 s when the 4B answers (about 70%); 2.85 s overall, slowest 5% 9.82 s. Per decision, not per task |
| Decision quality | JevBench v1.4.2, 231 public items | Brain 4B 0.835 · Brain 0.8B 0.723 · router 0.797; no sealed score yet |
| Visual grounding | ScreenSpot-Pro, 1,581 items, one pass | Eyes 4B 67.7% on a GPU (bf16, native resolution; base model 64.8%); 50.9% as the Mac app runs it (4-bit MLX, ≤ 2 MP) |
Our own runs. Methods and full tables: Brain results · Bench reference · Eyes results · Results and limits.
What we are working on next: ROADMAP.
Inference runs locally by default. The models download once, from Hugging Face or ModelScope; after that, deciding a step needs no network. A cloud model is opt-in: the app never calls one, but if you point the router's escalation tier at one yourself, the steps routed to it go to that service. Apps that DeskMind drives, such as a web page or a music app, still talk to their own servers.
area: …). Pull requests go to the repository that holds the code.
The frame from our logo, come to life. Artwork in brand/, rules in BRAND.md.
app/ and the CI): Apache-2.0; see NOTICE.Small enough to run on your Mac. Smart enough to ask. An open-source computer use agent for macOS, with local models (MLX) that see, decide and act. 得心,应手。
See the code
Small enough to run on your Mac. Smart enough to ask.
Open-source computer use for your Mac. DeskMind reads the screen, works out the next step and acts, all with small models running on your Mac. When a task could mean two things, it asks you instead of guessing.
中文 · Website · Docs · Download for Mac · Models · Discussions · Roadmap · Changelog
Watch the 56-second demo on deskmind.dev: a real recording with the released model. Two orders match "Lisa Wong", so it asks which one before writing.
How it was built, and what went wrong on the way: 3 weeks, 20 training rounds, $600. If DeskMind is useful or interesting to you, a star on this repo helps other people find it.
A small model, on your Mac. A 0.8B model decides each step and hands the unsure ones to a 4B. Both run on your Mac: no cloud round-trip, no per-step bill. The 0.8B decides in about 0.5 s; the 4B answers in about 3.6 s (median decision times).
System One: choices, not guesses. Each step is a multiple-choice question. The model scores every option instead of writing text, so every option gets a probability. Unsure steps go to the 4B or to you, and any agent can call it through POST /v1/systemone. In 39 real-desktop runs it never said "done" when the task was not done.
Open, from eyes to hands. Eyes, Brain, Hands and the Mac app are open source, along with Bench, which grades them. Every result below comes with its sample size, so you can reproduce it.
Use the Mac app. Download DeskMind for Mac (macOS 15+, Apple Silicon, signed and notarized). It ships the G18b release models. On first run it downloads the models (about 5.3 GB) and walks you through the permissions: Install the app.
Or run the model yourself. This runs the 4B alone (release G18b). It answers one step of a desktop task; it does not drive the desktop by itself.
# Terminal 1: get Brain and the model, then serve it
git clone https://github.com/deskmind-ai/brain && cd brain && uv sync --extra mlx
uv run hf download deskmind/brain-4b --revision g18b-q8 --local-dir models/brain-4b
uv run deskmind-brain-serve --predictor mlx:models/brain-4b --port 8793 --two-stage
The 4B is about 4.5 GB to download. If the download fails with a CAS Client Error, rerun it with HF_HUB_DISABLE_XET=1 in front. In mainland China, ModelScope carries the same files:
uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b.
# Terminal 2 (in the same brain folder): ask for the next step of a real desktop task
curl -s localhost:8793/v1/systemone -H 'Content-Type: application/json' -d @examples/request.json
Success looks like a typed decision with a probability for every option, e.g.
{"answers": {"operation": {"choice": "CLICK", "probabilities": {"CLICK": 0.96, "OPEN": 0.005, …}}, "click_target": {…}}}.
A probability is the model's weighting of the options, not a guarantee that the step is right.
The router (0.8B → 4B), as released: also download deskmind/brain-0.8b --revision g18b-q8 into models/brain-0.8b (0.8 GB), then serve
uv run deskmind-brain-serve --predictor mlx:models/brain-0.8b --escalate-to mlx:models/brain-4b --two-stage --port 8796.
The threshold (0.96) ships with the weights; each reply adds a routing record such as {"by": "strong", "reason": "low_conf", "fast_conf": 0.956}.
Step by step, with the full reply explained: Quickstart. To drive the desktop, add Hands.
| Role | ||
|---|---|---|
| Eyes | finds the target on screen | 4B visual grounder, for apps without an accessibility tree |
| Brain | decides the next step | 0.8B and 4B, MLX, 0.8B → 4B routing, /v1/systemone |
| Hands | observes and acts on macOS | accessibility and vision modes, budgets, cancellation |
| App (this repository) | brings it to your Mac | native app with a background helper; download |
| Bench | checks what really happened | sandbox desktop tasks with strict final-state graders |
Models: huggingface.co/deskmind (brain-0.8b, brain-4b, eyes-4b). Website: deskmind.dev. Docs: deskmind.dev/docs.
| Path | What |
|---|---|
app/ | The Mac app's source (Swift) and its build; build it yourself |
docs/ | How the app, its helper and the local model servers fit together |
brand/ | Logo, Xiaofang and social images; rules in BRAND.md |
.github/workflows/app.yml | CI: every app change is built and tested; a version tag is signed, notarized and drafted as a release |
ROADMAP.md | What we are working on next |
Mac app releases are published here: Releases; what changed in each one, in short: CHANGELOG.md.
| What | Setting | Result |
|---|---|---|
| Real-desktop tasks | Bench v25, 13 tasks × 3 runs, strict graders; router G18b (0.8B → 4B, 8-bit, threshold 0.96), through the app, one M4 Pro (48 GB) | 39/39 passed; 0 false "done" |
| Decision time | the same 39 runs, 208 decisions | median 0.48 s when the 0.8B answers (about 30% of steps), 3.6 s when the 4B answers (about 70%); 2.85 s overall, slowest 5% 9.82 s. Per decision, not per task |
| Decision quality | JevBench v1.4.2, 231 public items | Brain 4B 0.835 · Brain 0.8B 0.723 · router 0.797; no sealed score yet |
| Visual grounding | ScreenSpot-Pro, 1,581 items, one pass | Eyes 4B 67.7% on a GPU (bf16, native resolution; base model 64.8%); 50.9% as the Mac app runs it (4-bit MLX, ≤ 2 MP) |
Our own runs. Methods and full tables: Brain results · Bench reference · Eyes results · Results and limits.
What we are working on next: ROADMAP.
Inference runs locally by default. The models download once, from Hugging Face or ModelScope; after that, deciding a step needs no network. A cloud model is opt-in: the app never calls one, but if you point the router's escalation tier at one yourself, the steps routed to it go to that service. Apps that DeskMind drives, such as a web page or a music app, still talk to their own servers.
area: …). Pull requests go to the repository that holds the code.
The frame from our logo, come to life. Artwork in brand/, rules in BRAND.md.
app/ and the CI): Apache-2.0; see NOTICE.