Press a key, speak, text lands at your cursor. Press another, get a spoken answer. Local-first, lightweight voice dictation and assistant for Linux. φ
See the codePress a key and speak, and Fono types into any app, answers as a voice assistant,
drives your coding agent or smart home. It's a complete voice-AI stack
(speech-to-text, natural voices, a local LLM, wake word, speaker ID)
in one small binary. Everything runs locally, and every stage can switch
to a cloud provider if that fits you better.
Website · Docs · Install · Quickstart · Providers · Coding agents · Roadmap · Contributing
No setup marathon, no account, no stack of services to babysit. One small file: install it, press a key, talk. And it's not just for you; anything on your network can use it too: Home Assistant, Open WebUI, your editor. (If you're an engineer: think SQLite of voice AI. The entire stack, self-contained, in one binary.)
curl -fsSL https://fono.page/install | sh
The script picks the right binary for your CPU (and the Vulkan build if you have a GPU), starts the daemon, and opens the setup wizard in the same terminal. Everything runs locally unless you opt into a cloud provider.
Already installed? fono update keeps you current. Prefer packages? See other ways to install.
Two keys, one escape hatch. Each key auto-detects how you press it: a quick tap toggles recording, holding it turns into push-to-talk that stops on release.
| Key | What it does |
|---|---|
F7 | Dictate. Your words are typed into the focused window. Terminal, browser, editor, anything. |
F8 | Ask. Talk to an AI and get the reply read aloud, streamed sentence-by-sentence so audio starts before the model finishes thinking. |
Esc | Cancel a recording, or interrupt an assistant reply. |
F7 mic ▸ speech-to-text ▸ optional LLM polish ▸ typed into the focused window (+ clipboard)
F8 mic ▸ speech-to-text ▸ LLM assistant ▸ spoken reply, streamed into TTS
AI voice ▸ MCP ▸ Claude Code / Cursor / Forge ▸ the agent talks back and listens for more
Prefer no keys at all? Fono can idle and listen for a spoken wake phrase instead. Detection runs locally, and it's off until you enable it.
While you speak, a small overlay shows what the microphone hears — bars, oscilloscope, fft, or heatmap. Switch via the tray or [overlay].style in ~/.config/fono/config.toml.
fono update probes your host and pulls the matching build automatically.Local by default; every stage can be swapped to a cloud provider independently. Full matrix with models and config keys: docs/providers.md.
| Stage | Local (default) | Cloud |
|---|---|---|
| Speech-to-text | Whisper (bundled) | Groq · OpenAI · Gemini · Deepgram · Cartesia · AssemblyAI · Speechmatics · ElevenLabs |
| Polish | llama.cpp (bundled)* · Ollama | Cerebras · Groq · OpenAI · Anthropic · Gemini · OpenRouter |
| Assistant | llama.cpp (bundled)* · Ollama | OpenAI · Groq · Anthropic · Cerebras · Gemini · OpenRouter |
| Realtime assistant | — | Gemini Live (hands-free, back-and-forth conversation) |
| Text-to-speech | Kokoro (En) · Piper (International) · Supertonic (31 languages, opt-in) | OpenAI · Groq · Gemini · OpenRouter · Cartesia · Deepgram · ElevenLabs |
* Polish and the assistant share a single llama.cpp instance — the local model is loaded once, not twice.
Switching is one command or web config save. No restart necessary, the daemon hot-reloads:
fono use cloud groq # one key covers STT + polish + assistant + TTS
fono use stt deepgram # change a single stage
fono use tts cartesia
fono use local # back to fully local
fono keys add GROQ_API_KEY # keys live in ~/.config/fono/secrets.toml
fono keys check # reachability probe per stored key
Fono doesn't need a desktop. sudo fono install --server sets up a hardened systemd service on a headless box, and that one machine becomes the voice backend for everything else: Home Assistant discovers it over mDNS as a Wyoming speech-to-text, text-to-speech, and wake-word provider; editors, Open WebUI, llm, and LangChain talk to its OpenAI- and Ollama-compatible API on port 11434; and other Fono desktops on the LAN route their dictation through it. Audio stays on your network. There's a Docker container too — see docs/home-assistant.md.
Pick Settings… in the tray and every option opens as a searchable page in your browser. Saves apply instantly, no restart. It's a local page served by the daemon itself — loopback-only, off until you open it, and API keys are write-only: the page can set them but never read them back. Curious? Click around the demo.
Local-first, by design. With the default setup, audio and text never leave your machine. Cloud providers are strictly opt-in, per stage. The full data-flow map — what leaves, when, and to whom — is in docs/privacy.md.
.deb, .pkg.tar.zst, and .txz files are built by CI and attached to each release. They are not regularly tested — file an issue if one misbehaves.fono-vX.Y.Z-aarch64-apple-darwin binary. Download it, chmod +x, and run fono install — it sets up start-at-login and walks you through the one-time permission grants. It's only been tested on a headless remote Mac so far, not eyeballed on a real display yet — if you try it, an issue report (good or bad) is genuinely useful. Details in docs/build-macos.md.fono-vX.Y.Z-x86_64.exe. Download it and run fono install — it copies the app into your user folder and starts it at login, no administrator prompt. One download uses your GPU when a driver is present and falls back to the processor otherwise. This is an early port, built and exercised remotely rather than daily-driven, so expect rough edges — if you try it, an issue report (good or bad) is genuinely useful. Details in docs/build-windows.md.config.tomlLinux-first; used daily by the maintainer. macOS support is new and has not yet run on a real display — see Other ways to install. Rough edges exist — issues and patches are welcome. See the roadmap for what's next.
Pull requests welcome. See CONTRIBUTING.md for the workflow (DCO sign-off required).
GPL-3.0-only. See LICENSE.
Rust
92.2%
JavaScript
2.9%
Shell
2.2%
HTML
1.2%
Press a key, speak, text lands at your cursor. Press another, get a spoken answer. Local-first, lightweight voice dictation and assistant for Linux. φ
See the codePress a key and speak, and Fono types into any app, answers as a voice assistant,
drives your coding agent or smart home. It's a complete voice-AI stack
(speech-to-text, natural voices, a local LLM, wake word, speaker ID)
in one small binary. Everything runs locally, and every stage can switch
to a cloud provider if that fits you better.
Website · Docs · Install · Quickstart · Providers · Coding agents · Roadmap · Contributing
No setup marathon, no account, no stack of services to babysit. One small file: install it, press a key, talk. And it's not just for you; anything on your network can use it too: Home Assistant, Open WebUI, your editor. (If you're an engineer: think SQLite of voice AI. The entire stack, self-contained, in one binary.)
curl -fsSL https://fono.page/install | sh
The script picks the right binary for your CPU (and the Vulkan build if you have a GPU), starts the daemon, and opens the setup wizard in the same terminal. Everything runs locally unless you opt into a cloud provider.
Already installed? fono update keeps you current. Prefer packages? See other ways to install.
Two keys, one escape hatch. Each key auto-detects how you press it: a quick tap toggles recording, holding it turns into push-to-talk that stops on release.
| Key | What it does |
|---|---|
F7 | Dictate. Your words are typed into the focused window. Terminal, browser, editor, anything. |
F8 | Ask. Talk to an AI and get the reply read aloud, streamed sentence-by-sentence so audio starts before the model finishes thinking. |
Esc | Cancel a recording, or interrupt an assistant reply. |
F7 mic ▸ speech-to-text ▸ optional LLM polish ▸ typed into the focused window (+ clipboard)
F8 mic ▸ speech-to-text ▸ LLM assistant ▸ spoken reply, streamed into TTS
AI voice ▸ MCP ▸ Claude Code / Cursor / Forge ▸ the agent talks back and listens for more
Prefer no keys at all? Fono can idle and listen for a spoken wake phrase instead. Detection runs locally, and it's off until you enable it.
While you speak, a small overlay shows what the microphone hears — bars, oscilloscope, fft, or heatmap. Switch via the tray or [overlay].style in ~/.config/fono/config.toml.
fono update probes your host and pulls the matching build automatically.Local by default; every stage can be swapped to a cloud provider independently. Full matrix with models and config keys: docs/providers.md.
| Stage | Local (default) | Cloud |
|---|---|---|
| Speech-to-text | Whisper (bundled) | Groq · OpenAI · Gemini · Deepgram · Cartesia · AssemblyAI · Speechmatics · ElevenLabs |
| Polish | llama.cpp (bundled)* · Ollama | Cerebras · Groq · OpenAI · Anthropic · Gemini · OpenRouter |
| Assistant | llama.cpp (bundled)* · Ollama | OpenAI · Groq · Anthropic · Cerebras · Gemini · OpenRouter |
| Realtime assistant | — | Gemini Live (hands-free, back-and-forth conversation) |
| Text-to-speech | Kokoro (En) · Piper (International) · Supertonic (31 languages, opt-in) | OpenAI · Groq · Gemini · OpenRouter · Cartesia · Deepgram · ElevenLabs |
* Polish and the assistant share a single llama.cpp instance — the local model is loaded once, not twice.
Switching is one command or web config save. No restart necessary, the daemon hot-reloads:
fono use cloud groq # one key covers STT + polish + assistant + TTS
fono use stt deepgram # change a single stage
fono use tts cartesia
fono use local # back to fully local
fono keys add GROQ_API_KEY # keys live in ~/.config/fono/secrets.toml
fono keys check # reachability probe per stored key
Fono doesn't need a desktop. sudo fono install --server sets up a hardened systemd service on a headless box, and that one machine becomes the voice backend for everything else: Home Assistant discovers it over mDNS as a Wyoming speech-to-text, text-to-speech, and wake-word provider; editors, Open WebUI, llm, and LangChain talk to its OpenAI- and Ollama-compatible API on port 11434; and other Fono desktops on the LAN route their dictation through it. Audio stays on your network. There's a Docker container too — see docs/home-assistant.md.
Pick Settings… in the tray and every option opens as a searchable page in your browser. Saves apply instantly, no restart. It's a local page served by the daemon itself — loopback-only, off until you open it, and API keys are write-only: the page can set them but never read them back. Curious? Click around the demo.
Local-first, by design. With the default setup, audio and text never leave your machine. Cloud providers are strictly opt-in, per stage. The full data-flow map — what leaves, when, and to whom — is in docs/privacy.md.
.deb, .pkg.tar.zst, and .txz files are built by CI and attached to each release. They are not regularly tested — file an issue if one misbehaves.fono-vX.Y.Z-aarch64-apple-darwin binary. Download it, chmod +x, and run fono install — it sets up start-at-login and walks you through the one-time permission grants. It's only been tested on a headless remote Mac so far, not eyeballed on a real display yet — if you try it, an issue report (good or bad) is genuinely useful. Details in docs/build-macos.md.fono-vX.Y.Z-x86_64.exe. Download it and run fono install — it copies the app into your user folder and starts it at login, no administrator prompt. One download uses your GPU when a driver is present and falls back to the processor otherwise. This is an early port, built and exercised remotely rather than daily-driven, so expect rough edges — if you try it, an issue report (good or bad) is genuinely useful. Details in docs/build-windows.md.config.tomlLinux-first; used daily by the maintainer. macOS support is new and has not yet run on a real display — see Other ways to install. Rough edges exist — issues and patches are welcome. See the roadmap for what's next.
Pull requests welcome. See CONTRIBUTING.md for the workflow (DCO sign-off required).
GPL-3.0-only. See LICENSE.
Rust
92.2%
JavaScript
2.9%
Shell
2.2%
HTML
1.2%