rohanprichard/talktome

Voice calls with the coding agent you already use. A macOS menu-bar app.

Python

3

2 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Talktome - let your AI Agent call you

4

Sep 29, 2026

README

talktome

TalkToMe is a macOS menu-bar app for voice calls with a coding agent. The agent rings you from the session that it already runs. You answer, and you talk.

The agent keeps its model, tools, files, and history. TalkToMe does not supply a language model. The app records the microphone, turns speech into text, sends the text to the agent, and speaks the reply.

How a call works

  1. You ask the agent in its session to call you, for example "Call me with TalkToMe."
  2. The agent runs talktome call with its session ID.
  3. TalkToMe shows a ring on the screen. You select Answer.
  4. The agent speaks a short greeting.
  5. You speak. The agent receives your words as a message in its session and replies.
  6. You select End, or the agent runs talktome end.

The app has no window that starts a conversation. After setup, TalkToMe stays in the menu bar and waits for a call. The menu-bar icon opens Settings and gives controls to answer, decline, mute, and end a call.

Requirements

  • macOS. The app received tests only on macOS. The disk image is for Apple silicon (arm64).
  • Node.js 22 or later, for a build from source.
  • uv. uv installs Python 3.11 to 3.13 when necessary.
  • Xcode command-line tools, to build the disk image. The build compiles a small Swift helper.
  • One agent host: Codex, Claude Code, Hermes Agent, OpenClaw, or another host that can run shell commands.

Install and start

git clone https://github.com/rohanprichard/talktome.git
cd talktome
npm ci
uv sync --frozen
npm start

npm start runs uv sync --frozen if the Python environment is missing. Then it starts Electron.

First-run setup

The first start opens a setup window with six steps:

  1. Welcome. The window explains the call flow.
  2. Microphone. Select Allow microphone. macOS asks for permission.
  3. ElevenLabs. Enter an ElevenLabs API key, or select Later to use local speech.
  4. Agent connection. Select Install. This installs the agent skill and the talktome command.
  5. Glow color. Select the color that the call surface shows during a live call.
  6. All set. Ask your agent to call.

You can skip a step with Later. Settings contains the same options. To run setup again, quit the app and run npm run reset. This command clears the saved speech settings and the setup progress. It keeps the downloaded models and the local token.

Connect an agent

Install writes the skill file to ~/.codex/skills/talktome/SKILL.md. It also writes the skill to ~/.hermes/skills/ and ~/.openclaw/skills/ if those directories exist. It installs the talktome command in ~/.local/bin, /opt/homebrew/bin, or /usr/local/bin. The skill tells the agent which commands to run. See the skill.

For Claude Code, or for another host, copy the skill into the host's skill directory:

mkdir -p ~/.claude/skills/talktome
talktome skill > ~/.claude/skills/talktome/SKILL.md
HostCommandHow replies reach the call
Codextalktome call --agent codex --thread "$CODEX_THREAD_ID"TalkToMe reads the session's public replies automatically.
Claude Codetalktome call --agent claude --thread SESSION_IDThe agent runs talktome listen and talktome reply.
Hermes Agent chattalktome call --agent hermes --connection cooperative --thread IDThe agent runs talktome listen and talktome reply.
OpenClaw chattalktome call --agent openclaw --connection cooperative --thread IDThe agent runs talktome listen and talktome reply.
Other hoststalktome call --agent generic --thread IDThe agent runs talktome listen and talktome reply.

The listen and reply commands are the "cooperative" connection. They work with any host that can run shell commands on the Mac that runs TalkToMe. The commands exchange private files with the app. Thus, they work when a sandbox blocks local network access.

Hermes and OpenClaw also have experimental adapters for an API session or a Gateway session. Run talktome providers to see the connection methods that are ready. Agent support gives the setup and the limits.

Agent protocol gives the commands, the ring flow, and the local HTTP interface.

Remote bridge (experimental)

The remote bridge lets an agent on another server ring the laptop. Only text and call events cross the bridge. Microphone audio stays on the laptop. The setup uses SSH:

  1. Install talktome on the server: uv tool install git+https://github.com/rohanprichard/talktome.
  2. On the laptop, run talktome remote-connect user@server --install-service.
  3. Restart TalkToMe.

The app opens an SSH tunnel to the server itself, so SSH must log in with a key and no password prompt. This feature is experimental. See remote bridge.

During a call

After the agent finishes its reply, speak to start the next turn.

TalkToMe uses Smart Turn, a small local model, to decide when you finished speaking. Settings can select a fixed pause instead. See Smart Turn. Call latency and call timing describe the delays in a call.

Settings also sets the position of the call surface: Bottom or Top center.

Speech providers

FunctionLocal optionElevenLabs option
Speech recognitionWhisper Small (484 MB) or Whisper Base English (145 MB)Scribe v2
Agent voiceSystem voice or KokoroFlash v2.5

Whisper is the default for recognition. The system voice is the default voice. Download a Whisper model in Settings before you use local recognition. Kokoro downloads a 114 MB model and a 28 MB voice file. The app examines their SHA-256 hashes before use. Local models run offline after the download.

The ElevenLabs key needs access to the voice list and to each selected speech service. Select Remember key to keep the key in the macOS keychain. If you do not, the key stays in server memory until the app closes. The app never returns the key to the interface or writes it to its settings file. Remove key removes the key from the app session and from the keychain. See speech providers for the exact interfaces.

Data and network access

The server listens only on 127.0.0.1:8765. A generated local token protects its interface. The desktop windows use an HTTP-only session cookie. Other web origins cannot use the interface. The app keeps its token, settings, and models in ~/Library/Application Support/talktome. The app keeps up to 200 transcript messages and 512 events in memory. Closing the app clears them.

The app connects to the network for these purposes only:

  • Hugging Face, to download Whisper and the Smart Turn model.
  • GitHub, to download the Kokoro model and voice file.
  • ElevenLabs, only if you select an ElevenLabs service. ElevenLabs recognition sends microphone audio. ElevenLabs voice sends reply text. Service charges and the provider's retention rules apply.
  • A Hermes or OpenClaw host, or a relay, only if you configure one.

The agent receives the text of what you say. The agent's provider and tools have their own data rules. The cooperative command files contain conversation text. Settings for Hermes and OpenClaw are in agent-hosts.json in the data directory. This file holds a plaintext token with mode 0600.

Environment settings:

NamePurpose
TALKTOME_DATA_DIRChange the local data directory
TALKTOME_PORTChange the desktop server port
TALKTOME_URLSet the server address for external clients
TALKTOME_TOKENSupply an existing shared token
TALKTOME_RELOADRestart the server when Python files change. Development only.
TALKTOME_FLOATING_CALLSet to 0 to turn off the call window. The call then has no controls on screen.
TALKTOME_ALLOW_REMOTE_AGENTSSet to 1 to permit a Hermes host that is not on this Mac. The URL must use HTTPS.

Build the disk image

npm run build:app

This command freezes the Python server into one binary, draws the icon, compiles the notch helper, and runs electron-builder. The result is dist/app/TalkToMe-<version>-arm64.dmg. The script mounts the disk image after the build and examines its contents. To reuse the last frozen server when only the desktop code changed, run npm run build:dmg.

To install the app, open the disk image and drag TalkToMe to Applications. The build is not signed. It opens on the Mac that built it. Other Macs block it, because notarization needs a paid Apple Developer ID. The bundle does not include the speech models. The app downloads them at first use.

Development

uv sync --frozen        # install the Python environment
npm ci                  # install the Node packages
npm start               # start the app from source
uv run pytest -q        # Python tests
uv run ruff check       # Python lint
node --test tests/      # JavaScript tests

npm test runs the Python tests and the JavaScript tests together. These commands start the real app for end-to-end checks:

CommandWhat it examines
npm run test:callThe call window: position, stacking, and growth of the transcript
npm run test:attach-callA full call against a real Codex session. It sends a few short model requests.
npm run test:streamTime to first audio for ElevenLabs. It spends credits on two short replies.

The call test turns off the Chromium sandbox. A normal start keeps the sandbox on.

The development notes hold plans, research, and a work log. They can be out of date.

Security

To report a security problem, read SECURITY.md.

Contributing

Read CONTRIBUTING.md before you open a pull request.

License

MIT. See LICENSE. TalkToMe uses Electron, FastAPI, faster-whisper, kokoro-onnx, Pipecat Smart Turn, and optional ElevenLabs services. It does not contain copied SpeakType or AgentCall code. NOTICE lists the third-party references.

claude-code
codex
coding-agents
hermes-agent
macos
speech-to-text
voice-assistant

rohanprichard/talktome

Voice calls with the coding agent you already use. A macOS menu-bar app.

Python

3

2 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Talktome - let your AI Agent call you

4

Sep 29, 2026

README

talktome

TalkToMe is a macOS menu-bar app for voice calls with a coding agent. The agent rings you from the session that it already runs. You answer, and you talk.

The agent keeps its model, tools, files, and history. TalkToMe does not supply a language model. The app records the microphone, turns speech into text, sends the text to the agent, and speaks the reply.

How a call works

  1. You ask the agent in its session to call you, for example "Call me with TalkToMe."
  2. The agent runs talktome call with its session ID.
  3. TalkToMe shows a ring on the screen. You select Answer.
  4. The agent speaks a short greeting.
  5. You speak. The agent receives your words as a message in its session and replies.
  6. You select End, or the agent runs talktome end.

The app has no window that starts a conversation. After setup, TalkToMe stays in the menu bar and waits for a call. The menu-bar icon opens Settings and gives controls to answer, decline, mute, and end a call.

Requirements

  • macOS. The app received tests only on macOS. The disk image is for Apple silicon (arm64).
  • Node.js 22 or later, for a build from source.
  • uv. uv installs Python 3.11 to 3.13 when necessary.
  • Xcode command-line tools, to build the disk image. The build compiles a small Swift helper.
  • One agent host: Codex, Claude Code, Hermes Agent, OpenClaw, or another host that can run shell commands.

Install and start

git clone https://github.com/rohanprichard/talktome.git
cd talktome
npm ci
uv sync --frozen
npm start

npm start runs uv sync --frozen if the Python environment is missing. Then it starts Electron.

First-run setup

The first start opens a setup window with six steps:

  1. Welcome. The window explains the call flow.
  2. Microphone. Select Allow microphone. macOS asks for permission.
  3. ElevenLabs. Enter an ElevenLabs API key, or select Later to use local speech.
  4. Agent connection. Select Install. This installs the agent skill and the talktome command.
  5. Glow color. Select the color that the call surface shows during a live call.
  6. All set. Ask your agent to call.

You can skip a step with Later. Settings contains the same options. To run setup again, quit the app and run npm run reset. This command clears the saved speech settings and the setup progress. It keeps the downloaded models and the local token.

Connect an agent

Install writes the skill file to ~/.codex/skills/talktome/SKILL.md. It also writes the skill to ~/.hermes/skills/ and ~/.openclaw/skills/ if those directories exist. It installs the talktome command in ~/.local/bin, /opt/homebrew/bin, or /usr/local/bin. The skill tells the agent which commands to run. See the skill.

For Claude Code, or for another host, copy the skill into the host's skill directory:

mkdir -p ~/.claude/skills/talktome
talktome skill > ~/.claude/skills/talktome/SKILL.md
HostCommandHow replies reach the call
Codextalktome call --agent codex --thread "$CODEX_THREAD_ID"TalkToMe reads the session's public replies automatically.
Claude Codetalktome call --agent claude --thread SESSION_IDThe agent runs talktome listen and talktome reply.
Hermes Agent chattalktome call --agent hermes --connection cooperative --thread IDThe agent runs talktome listen and talktome reply.
OpenClaw chattalktome call --agent openclaw --connection cooperative --thread IDThe agent runs talktome listen and talktome reply.
Other hoststalktome call --agent generic --thread IDThe agent runs talktome listen and talktome reply.

The listen and reply commands are the "cooperative" connection. They work with any host that can run shell commands on the Mac that runs TalkToMe. The commands exchange private files with the app. Thus, they work when a sandbox blocks local network access.

Hermes and OpenClaw also have experimental adapters for an API session or a Gateway session. Run talktome providers to see the connection methods that are ready. Agent support gives the setup and the limits.

Agent protocol gives the commands, the ring flow, and the local HTTP interface.

Remote bridge (experimental)

The remote bridge lets an agent on another server ring the laptop. Only text and call events cross the bridge. Microphone audio stays on the laptop. The setup uses SSH:

  1. Install talktome on the server: uv tool install git+https://github.com/rohanprichard/talktome.
  2. On the laptop, run talktome remote-connect user@server --install-service.
  3. Restart TalkToMe.

The app opens an SSH tunnel to the server itself, so SSH must log in with a key and no password prompt. This feature is experimental. See remote bridge.

During a call

After the agent finishes its reply, speak to start the next turn.

TalkToMe uses Smart Turn, a small local model, to decide when you finished speaking. Settings can select a fixed pause instead. See Smart Turn. Call latency and call timing describe the delays in a call.

Settings also sets the position of the call surface: Bottom or Top center.

Speech providers

FunctionLocal optionElevenLabs option
Speech recognitionWhisper Small (484 MB) or Whisper Base English (145 MB)Scribe v2
Agent voiceSystem voice or KokoroFlash v2.5

Whisper is the default for recognition. The system voice is the default voice. Download a Whisper model in Settings before you use local recognition. Kokoro downloads a 114 MB model and a 28 MB voice file. The app examines their SHA-256 hashes before use. Local models run offline after the download.

The ElevenLabs key needs access to the voice list and to each selected speech service. Select Remember key to keep the key in the macOS keychain. If you do not, the key stays in server memory until the app closes. The app never returns the key to the interface or writes it to its settings file. Remove key removes the key from the app session and from the keychain. See speech providers for the exact interfaces.

Data and network access

The server listens only on 127.0.0.1:8765. A generated local token protects its interface. The desktop windows use an HTTP-only session cookie. Other web origins cannot use the interface. The app keeps its token, settings, and models in ~/Library/Application Support/talktome. The app keeps up to 200 transcript messages and 512 events in memory. Closing the app clears them.

The app connects to the network for these purposes only:

  • Hugging Face, to download Whisper and the Smart Turn model.
  • GitHub, to download the Kokoro model and voice file.
  • ElevenLabs, only if you select an ElevenLabs service. ElevenLabs recognition sends microphone audio. ElevenLabs voice sends reply text. Service charges and the provider's retention rules apply.
  • A Hermes or OpenClaw host, or a relay, only if you configure one.

The agent receives the text of what you say. The agent's provider and tools have their own data rules. The cooperative command files contain conversation text. Settings for Hermes and OpenClaw are in agent-hosts.json in the data directory. This file holds a plaintext token with mode 0600.

Environment settings:

NamePurpose
TALKTOME_DATA_DIRChange the local data directory
TALKTOME_PORTChange the desktop server port
TALKTOME_URLSet the server address for external clients
TALKTOME_TOKENSupply an existing shared token
TALKTOME_RELOADRestart the server when Python files change. Development only.
TALKTOME_FLOATING_CALLSet to 0 to turn off the call window. The call then has no controls on screen.
TALKTOME_ALLOW_REMOTE_AGENTSSet to 1 to permit a Hermes host that is not on this Mac. The URL must use HTTPS.

Build the disk image

npm run build:app

This command freezes the Python server into one binary, draws the icon, compiles the notch helper, and runs electron-builder. The result is dist/app/TalkToMe-<version>-arm64.dmg. The script mounts the disk image after the build and examines its contents. To reuse the last frozen server when only the desktop code changed, run npm run build:dmg.

To install the app, open the disk image and drag TalkToMe to Applications. The build is not signed. It opens on the Mac that built it. Other Macs block it, because notarization needs a paid Apple Developer ID. The bundle does not include the speech models. The app downloads them at first use.

Development

uv sync --frozen        # install the Python environment
npm ci                  # install the Node packages
npm start               # start the app from source
uv run pytest -q        # Python tests
uv run ruff check       # Python lint
node --test tests/      # JavaScript tests

npm test runs the Python tests and the JavaScript tests together. These commands start the real app for end-to-end checks:

CommandWhat it examines
npm run test:callThe call window: position, stacking, and growth of the transcript
npm run test:attach-callA full call against a real Codex session. It sends a few short model requests.
npm run test:streamTime to first audio for ElevenLabs. It spends credits on two short replies.

The call test turns off the Chromium sandbox. A normal start keeps the sandbox on.

The development notes hold plans, research, and a work log. They can be out of date.

Security

To report a security problem, read SECURITY.md.

Contributing

Read CONTRIBUTING.md before you open a pull request.

License

MIT. See LICENSE. TalkToMe uses Electron, FastAPI, faster-whisper, kokoro-onnx, Pipecat Smart Turn, and optional ElevenLabs services. It does not contain copied SpeakType or AgentCall code. NOTICE lists the third-party references.

claude-code
codex
coding-agents
hermes-agent
macos
speech-to-text
voice-assistant

Languages

Python

62.5%

JavaScript

25.1%

CSS

4.7%

HTML

4.0%

Swift

3.1%