Autobot is a privacy-aware operating harness that integrates via native ChatGPT Project capabilities integrating OpenAI synchronous voice and OAI UI and control planes. This is an internal, knowledge worker tool leveraged by the team at Autonomous Production tool
See the codeAutoBot: a self-improving agentic harness that makes frontier AI better at finishing complex knowledge work.
AutoBot achieved 18.5% higher task completion than the published OpenAI Sol Max baseline, surpassing Anthropic’s Claude Opus 5 Max on OSWorld 2.0, a benchmark of long, multi-application workflows. It also reached #1 on the official AssistantBench hidden-test leaderboard. Results and methodology
Hard workflows become upgrades to the agent itself. AutoBot repairs its own harness, independently validates the changes, and carries them forward. The next workflow inherits the improvement. Compounding capability, without retraining the model.
Your knowledge outgrows the context window. Hierarchical memory lives on disk; task-specific retrieval builds the working context. Nightly consolidation integrates new knowledge and corrections. Your agent accumulates institutional memory across projects and conversations.
Your project can outlive the agent working on it. Persistent task graphs, atomic checkpoints and independent supervision let a replacement worker resume the assignment. Completion is bound to current requirements and verified destination evidence.
Local compute makes persistent intelligence economical. Your CPU handles orchestration, state and integrity checks. Compiled context and reusable proofs reduce repeated inference, directing the model’s budget toward the difficult judgments that move work forward.
Open source. Native ChatGPT on your Mac. Built by Autonomous Production. Get AutoBot.
#1 on the recorded official hidden-test leaderboard: 50.70% accuracy across 181 tasks.
AutoAssist was the harness's previous brand name, changed to AutoBot two weeks after this screenshot.
32.41% binary accuracy and 64.28% partial accuracy across 108 tasks. Final best-valid-per-task aggregate, compared with the published September 10, 2026 leaderboard snapshot.
SHARED automatically.Autobot is the combination of two layers:
AGENTS.md, privacy zones, selective memory, project status, a local objective state machine, a one-minute liveness supervisor, external-action policy, and first-time setup.Autobot is not a separate model, chatbot, or agent gateway. It depends on current ChatGPT capabilities and their plan, region, usage, sandbox, and permission limits. See Architecture and Permissions.
This comparison evaluates Autobot plus native ChatGPT for one person managing personal and work tasks on a Mac. "Better" means a more explicit default for the stated user need, based on the published design; it does not mean a measured advantage in accuracy, speed, or reliability. "Doesn't do" means the specific built-in requirement was not found in the primary documentation reviewed on September 10, 2026. Both alternatives can be extended.
| Top 3 shared capabilities: Autobot's approach and user benefits | Top 3 additional Autobot capabilities and user benefits | |
|---|---|---|
| OpenClaw | 1. Remember context with explicit boundaries. OpenClaw persists and searches memory. Autobot adds default personal, family/friends, work, and opt-in shared zones. This makes the rules for reusing private context in a work task more explicit. Memory / Autobot privacy. 2. Track the promised outcome. OpenClaw records background tasks and delivery state. Autobot assigns each promised output an owner, destination, and completion gate. This gives the user a clearer record of what remains owed after an interruption. Tasks / Autobot contract. 3. Check the result of an app action. OpenClaw supports tools and outbound audit history. Autobot requires account, destination, and content checks before a write, followed by rendered inspection. This adds an explicit check for a wrong destination, failed save, or duplicate action. Audit / Autobot writes. | 1. Require a separate completion validator. Autobot's five-stage process requires current evidence and a validator label different from the producer, with destination readback for external work. This gives users evidence beyond the worker's own success report. Completion. 2. Block writes when clean delivery cannot be verified. Autobot's default policy requires first-party Computer Use and stops a route that forces unwanted attribution. This gives users an explicit publication rule instead of relying on each connector's behavior. Write policy. 3. Keep learned tone separate by channel and audience. Autobot requires authorized examples and compact, separate communication profiles. This helps keep a family-text style from shaping a customer email. Profiles. |
| Hermes | 1. Remember context with purpose-specific retrieval. Hermes has curated memory, user profiles, and session search. Autobot makes relationship and purpose part of its default retrieval rules. This gives users clearer control over which personal facts may inform professional work. Memory / Autobot privacy. 2. Follow work through to its deliverable. Hermes schedules jobs and delivers their outputs. Autobot also retains the objective, ordered evidence stages, and unfinished outputs. This makes it easier to distinguish a job that ran from a requested artifact that was saved and checked. Scheduling / Autobot completion. 3. Verify user-visible changes. Hermes provides tool approvals and execution controls. Autobot adds a standard before-and-after inspection of the signed-in destination. This gives the user a specific check that an authorized action produced the intended visible result. Security / Autobot writes. | 1. Require independent acceptance of every completion stage. Autobot separates the producer and validator labels and requires fresh destination evidence for external outputs. This makes an unsupported "done" report insufficient under the operating contract. Completion. 2. Require a clean first-party write route by default. Autobot checks the exact account and visible content, then rejects forced-attribution or unverifiable delivery routes. This gives users a consistent rule for communications across services. Write policy. 3. Require audience-specific tone-learning boundaries. Hermes supports personality and user-style preferences; Autobot specifies separately authorized profiles for each channel and audience. This gives users more explicit control over which examples shape each kind of message. Hermes memory / Autobot profiles. |
These are operating-contract differences, not guarantees of error-free execution. Autobot's privacy zones and validator separation are procedural within one Mac; they are not OS isolation or cryptographic identities. OpenClaw and Hermes offer broader standalone deployment, messaging, and model choices. The absence findings above concern the exact required workflows, not an absence of memory, privacy controls, voice, verification tools, or automation in either project.
The project-local first-time skill now acts as the setup operator. It checks current state, applies safe local defaults, runs supported actions, preserves the earliest unresolved user or capability gate, and returns to the user's original task only after current marker and status readback.
The base folder installs without Node or sign-in. The advanced runtime needs Node 22 or newer. Native app capabilities are discovered separately; no model, account, notification recipient or external write is configured for you. See capabilities and availability.
| Zone | Intended context | Default boundary |
|---|---|---|
PRIVATE | Sensitive personal facts, preferences, health, finances, and private plans | Never shared automatically |
FAMILY_FRIENDS | Relationships, events, logistics, and authorized personal communication patterns | Never used for work without a current explicit need |
WORK | Organizations, projects, teammates, customers, vendors, and professional communication patterns | Never used for personal communication unless explicitly relevant |
SHARED | The minimum facts you deliberately make reusable across zones | Opt-in only, with provenance |
These are owner-only folders and operating rules inside one macOS account. They are not separate encrypted vaults, macOS users, or hardware security boundaries. Use separate macOS accounts or separate machines when the trust domains require stronger isolation. Read Privacy.
Each durable objective moves through five ordered stages:
research_completedraft_completedestination_updatedsave_confirmedrendered_readback_verifiedThe local runtime enforces stage order, evidence hashes, and a validator label different from the producer label. The operating contract requires a genuinely separate validator to use fresh evidence and inspect the authoritative destination for external work.
This is procedural independence on one Mac, not a separate security enclave. Autobot cannot bypass a login, MFA, macOS permission, plan limit, user decision, or the scope of the user's instruction. The supervisor records liveness and flags stalled objectives. It does not control desktop apps in the background or manufacture permission to continue.
Autobot treats connectors, apps, MCP tools, browser integrations, and APIs as read-only unless a destination-specific adapter proves the same clean-write properties. Recipient-visible writes normally use the signed-in first-party interface through native ChatGPT Computer Use.
Before a live write, the operating contract requires the active ChatGPT workflow to check the app, account, destination, scope, and final visible content. Afterward, the workflow must inspect the rendered result and check for the intended mutation, no duplicate, no failure, and no added AI or ChatGPT attribution. If a platform forces a non-removable label or the result is ambiguous, the workflow must not claim success.
That policy cannot remove immutable metadata or disclosures controlled by a third-party platform. Details: Security.
ChatGPT Voice can start separate threads for longer tasks, check them, send follow-ups, and bring progress or blockers back into the voice conversation. Goal mode keeps the outcome and completion criteria attached to the work.
Important limits:
./runtime/bin/autoassist doctor
./runtime/bin/autoassist version
./runtime/bin/autoassist core status
./runtime/bin/autoassist core help
Run ./runtime/bin/autoassist help for the complete local command list.
privacy-scan is a clean release-candidate check, not a post-install scan of a populated user workspace. Run it only before user configuration or against a separate clean release tree.
Autobot 0.3.0 is an early public release for a single user on a Mac. The package is not a multi-tenant service, an OS sandbox, a cryptographic privacy boundary, or a guarantee that every third-party action will succeed. Treat a downloadable ZIP as verified only when its checksum and manifest match the adjacent release artifacts and the packaged candidate passes the bundled release checks.
Autobot is available under the MIT License. Copyright (c) 2026 Autobot contributors.
Set up AutoBot on your MacBook, then use it from your phone.
On your MacBook, download OpenAI's desktop app, sign in, and select Codex. The download is called ChatGPT, with Codex inside it.
AutoBot is free, but regular use will likely need a paid ChatGPT plan; choose one that fits your needs.
In Plugins > Computer Use, choose Install or Enable and turn on its server and skill switches. In your Mac's System Settings > Privacy & Security, enable the named OpenAI app or helper for:
Stay at your MacBook and plan for at least three permission-related app restarts during first setup; the number varies. Reopen the app after each prompt and continue.
Open Settings with Command + comma and choose these preferences:
Open the AutoBot release and download AutoAssist-v0.3.0.zip plus AutoAssist-v0.3.0.zip.sha256. Keep them together and double-click the ZIP to open the download.
Create a local project called AutoBot in Codex, choose the downloaded AutoAssist folder, and start a task there. Paste: “Install AutoBot from https://github.com/demeyer1/Autobot using INSTALL_FOR_AI.md, including required support software from official sources, then guide setup.”
After installation, choose Edit project > Add folder, select AutoAssist in your Mac's home folder, and choose Make primary. Start a fresh task there: “Run first-time setup for this installed folder.”
Install the ChatGPT mobile app, sign in to the same account and workspace, and on your Mac open Settings > Connections > Control this Mac or PC > Set up. Scan the QR code, finish verification, and open Remote on your phone.
Open Settings > Computer Use > Chrome > Install and add the ChatGPT extension to your chosen Chrome profile. Return to settings and check for Manage.
Tell AutoBot, “Connect and test my email, calendar, Slack, and iMessage,” then sign in to your chosen accounts, including Messages on your Mac. Use available plugins for reading and Computer Use in the signed-in apps for sending.
On your phone, open Remote > AutoBot and ask, “Make and save tomorrow's to-do list.” Open the result, then try voice.
Setup references: Mac permissions, permission choices, phone pairing, and MFA guidance.
If the user asks you to install this repository, read INSTALL_FOR_AI.md first. It contains the commands, local project setup, verification and resume instructions.
The user can start with: “Go to https://github.com/demeyer1/Autobot and install AutoBot in a new local project, following INSTALL_FOR_AI.md.”
65 commits
Shell
63.2%
JavaScript
35.8%
Python
1.1%
Autobot is a privacy-aware operating harness that integrates via native ChatGPT Project capabilities integrating OpenAI synchronous voice and OAI UI and control planes. This is an internal, knowledge worker tool leveraged by the team at Autonomous Production tool
See the codeAutoBot: a self-improving agentic harness that makes frontier AI better at finishing complex knowledge work.
AutoBot achieved 18.5% higher task completion than the published OpenAI Sol Max baseline, surpassing Anthropic’s Claude Opus 5 Max on OSWorld 2.0, a benchmark of long, multi-application workflows. It also reached #1 on the official AssistantBench hidden-test leaderboard. Results and methodology
Hard workflows become upgrades to the agent itself. AutoBot repairs its own harness, independently validates the changes, and carries them forward. The next workflow inherits the improvement. Compounding capability, without retraining the model.
Your knowledge outgrows the context window. Hierarchical memory lives on disk; task-specific retrieval builds the working context. Nightly consolidation integrates new knowledge and corrections. Your agent accumulates institutional memory across projects and conversations.
Your project can outlive the agent working on it. Persistent task graphs, atomic checkpoints and independent supervision let a replacement worker resume the assignment. Completion is bound to current requirements and verified destination evidence.
Local compute makes persistent intelligence economical. Your CPU handles orchestration, state and integrity checks. Compiled context and reusable proofs reduce repeated inference, directing the model’s budget toward the difficult judgments that move work forward.
Open source. Native ChatGPT on your Mac. Built by Autonomous Production. Get AutoBot.
#1 on the recorded official hidden-test leaderboard: 50.70% accuracy across 181 tasks.
AutoAssist was the harness's previous brand name, changed to AutoBot two weeks after this screenshot.
32.41% binary accuracy and 64.28% partial accuracy across 108 tasks. Final best-valid-per-task aggregate, compared with the published September 10, 2026 leaderboard snapshot.
SHARED automatically.Autobot is the combination of two layers:
AGENTS.md, privacy zones, selective memory, project status, a local objective state machine, a one-minute liveness supervisor, external-action policy, and first-time setup.Autobot is not a separate model, chatbot, or agent gateway. It depends on current ChatGPT capabilities and their plan, region, usage, sandbox, and permission limits. See Architecture and Permissions.
This comparison evaluates Autobot plus native ChatGPT for one person managing personal and work tasks on a Mac. "Better" means a more explicit default for the stated user need, based on the published design; it does not mean a measured advantage in accuracy, speed, or reliability. "Doesn't do" means the specific built-in requirement was not found in the primary documentation reviewed on September 10, 2026. Both alternatives can be extended.
| Top 3 shared capabilities: Autobot's approach and user benefits | Top 3 additional Autobot capabilities and user benefits | |
|---|---|---|
| OpenClaw | 1. Remember context with explicit boundaries. OpenClaw persists and searches memory. Autobot adds default personal, family/friends, work, and opt-in shared zones. This makes the rules for reusing private context in a work task more explicit. Memory / Autobot privacy. 2. Track the promised outcome. OpenClaw records background tasks and delivery state. Autobot assigns each promised output an owner, destination, and completion gate. This gives the user a clearer record of what remains owed after an interruption. Tasks / Autobot contract. 3. Check the result of an app action. OpenClaw supports tools and outbound audit history. Autobot requires account, destination, and content checks before a write, followed by rendered inspection. This adds an explicit check for a wrong destination, failed save, or duplicate action. Audit / Autobot writes. | 1. Require a separate completion validator. Autobot's five-stage process requires current evidence and a validator label different from the producer, with destination readback for external work. This gives users evidence beyond the worker's own success report. Completion. 2. Block writes when clean delivery cannot be verified. Autobot's default policy requires first-party Computer Use and stops a route that forces unwanted attribution. This gives users an explicit publication rule instead of relying on each connector's behavior. Write policy. 3. Keep learned tone separate by channel and audience. Autobot requires authorized examples and compact, separate communication profiles. This helps keep a family-text style from shaping a customer email. Profiles. |
| Hermes | 1. Remember context with purpose-specific retrieval. Hermes has curated memory, user profiles, and session search. Autobot makes relationship and purpose part of its default retrieval rules. This gives users clearer control over which personal facts may inform professional work. Memory / Autobot privacy. 2. Follow work through to its deliverable. Hermes schedules jobs and delivers their outputs. Autobot also retains the objective, ordered evidence stages, and unfinished outputs. This makes it easier to distinguish a job that ran from a requested artifact that was saved and checked. Scheduling / Autobot completion. 3. Verify user-visible changes. Hermes provides tool approvals and execution controls. Autobot adds a standard before-and-after inspection of the signed-in destination. This gives the user a specific check that an authorized action produced the intended visible result. Security / Autobot writes. | 1. Require independent acceptance of every completion stage. Autobot separates the producer and validator labels and requires fresh destination evidence for external outputs. This makes an unsupported "done" report insufficient under the operating contract. Completion. 2. Require a clean first-party write route by default. Autobot checks the exact account and visible content, then rejects forced-attribution or unverifiable delivery routes. This gives users a consistent rule for communications across services. Write policy. 3. Require audience-specific tone-learning boundaries. Hermes supports personality and user-style preferences; Autobot specifies separately authorized profiles for each channel and audience. This gives users more explicit control over which examples shape each kind of message. Hermes memory / Autobot profiles. |
These are operating-contract differences, not guarantees of error-free execution. Autobot's privacy zones and validator separation are procedural within one Mac; they are not OS isolation or cryptographic identities. OpenClaw and Hermes offer broader standalone deployment, messaging, and model choices. The absence findings above concern the exact required workflows, not an absence of memory, privacy controls, voice, verification tools, or automation in either project.
The project-local first-time skill now acts as the setup operator. It checks current state, applies safe local defaults, runs supported actions, preserves the earliest unresolved user or capability gate, and returns to the user's original task only after current marker and status readback.
The base folder installs without Node or sign-in. The advanced runtime needs Node 22 or newer. Native app capabilities are discovered separately; no model, account, notification recipient or external write is configured for you. See capabilities and availability.
| Zone | Intended context | Default boundary |
|---|---|---|
PRIVATE | Sensitive personal facts, preferences, health, finances, and private plans | Never shared automatically |
FAMILY_FRIENDS | Relationships, events, logistics, and authorized personal communication patterns | Never used for work without a current explicit need |
WORK | Organizations, projects, teammates, customers, vendors, and professional communication patterns | Never used for personal communication unless explicitly relevant |
SHARED | The minimum facts you deliberately make reusable across zones | Opt-in only, with provenance |
These are owner-only folders and operating rules inside one macOS account. They are not separate encrypted vaults, macOS users, or hardware security boundaries. Use separate macOS accounts or separate machines when the trust domains require stronger isolation. Read Privacy.
Each durable objective moves through five ordered stages:
research_completedraft_completedestination_updatedsave_confirmedrendered_readback_verifiedThe local runtime enforces stage order, evidence hashes, and a validator label different from the producer label. The operating contract requires a genuinely separate validator to use fresh evidence and inspect the authoritative destination for external work.
This is procedural independence on one Mac, not a separate security enclave. Autobot cannot bypass a login, MFA, macOS permission, plan limit, user decision, or the scope of the user's instruction. The supervisor records liveness and flags stalled objectives. It does not control desktop apps in the background or manufacture permission to continue.
Autobot treats connectors, apps, MCP tools, browser integrations, and APIs as read-only unless a destination-specific adapter proves the same clean-write properties. Recipient-visible writes normally use the signed-in first-party interface through native ChatGPT Computer Use.
Before a live write, the operating contract requires the active ChatGPT workflow to check the app, account, destination, scope, and final visible content. Afterward, the workflow must inspect the rendered result and check for the intended mutation, no duplicate, no failure, and no added AI or ChatGPT attribution. If a platform forces a non-removable label or the result is ambiguous, the workflow must not claim success.
That policy cannot remove immutable metadata or disclosures controlled by a third-party platform. Details: Security.
ChatGPT Voice can start separate threads for longer tasks, check them, send follow-ups, and bring progress or blockers back into the voice conversation. Goal mode keeps the outcome and completion criteria attached to the work.
Important limits:
./runtime/bin/autoassist doctor
./runtime/bin/autoassist version
./runtime/bin/autoassist core status
./runtime/bin/autoassist core help
Run ./runtime/bin/autoassist help for the complete local command list.
privacy-scan is a clean release-candidate check, not a post-install scan of a populated user workspace. Run it only before user configuration or against a separate clean release tree.
Autobot 0.3.0 is an early public release for a single user on a Mac. The package is not a multi-tenant service, an OS sandbox, a cryptographic privacy boundary, or a guarantee that every third-party action will succeed. Treat a downloadable ZIP as verified only when its checksum and manifest match the adjacent release artifacts and the packaged candidate passes the bundled release checks.
Autobot is available under the MIT License. Copyright (c) 2026 Autobot contributors.
Set up AutoBot on your MacBook, then use it from your phone.
On your MacBook, download OpenAI's desktop app, sign in, and select Codex. The download is called ChatGPT, with Codex inside it.
AutoBot is free, but regular use will likely need a paid ChatGPT plan; choose one that fits your needs.
In Plugins > Computer Use, choose Install or Enable and turn on its server and skill switches. In your Mac's System Settings > Privacy & Security, enable the named OpenAI app or helper for:
Stay at your MacBook and plan for at least three permission-related app restarts during first setup; the number varies. Reopen the app after each prompt and continue.
Open Settings with Command + comma and choose these preferences:
Open the AutoBot release and download AutoAssist-v0.3.0.zip plus AutoAssist-v0.3.0.zip.sha256. Keep them together and double-click the ZIP to open the download.
Create a local project called AutoBot in Codex, choose the downloaded AutoAssist folder, and start a task there. Paste: “Install AutoBot from https://github.com/demeyer1/Autobot using INSTALL_FOR_AI.md, including required support software from official sources, then guide setup.”
After installation, choose Edit project > Add folder, select AutoAssist in your Mac's home folder, and choose Make primary. Start a fresh task there: “Run first-time setup for this installed folder.”
Install the ChatGPT mobile app, sign in to the same account and workspace, and on your Mac open Settings > Connections > Control this Mac or PC > Set up. Scan the QR code, finish verification, and open Remote on your phone.
Open Settings > Computer Use > Chrome > Install and add the ChatGPT extension to your chosen Chrome profile. Return to settings and check for Manage.
Tell AutoBot, “Connect and test my email, calendar, Slack, and iMessage,” then sign in to your chosen accounts, including Messages on your Mac. Use available plugins for reading and Computer Use in the signed-in apps for sending.
On your phone, open Remote > AutoBot and ask, “Make and save tomorrow's to-do list.” Open the result, then try voice.
Setup references: Mac permissions, permission choices, phone pairing, and MFA guidance.
If the user asks you to install this repository, read INSTALL_FOR_AI.md first. It contains the commands, local project setup, verification and resume instructions.
The user can start with: “Go to https://github.com/demeyer1/Autobot and install AutoBot in a new local project, following INSTALL_FOR_AI.md.”
65 commits
Shell
63.2%
JavaScript
35.8%
Python
1.1%