Voice typing for Linux Mint XFCE. Double-tap Right Alt, speak, and clean formatted text is instantly pasted into whatever window you are using.
AutoType is a fast, flexible, open-source Linux alternative to Wispr Flow. It runs as a lightweight, silent background daemon: listening for a single hotkey gesture, converting speech to text, intelligently cleaning the transcription, and typing it into your active app.
Bring your own providers: AutoType is completely provider-agnostic. Choose any popular AI provider for text cleaning (OpenAI, Anthropic Claude, xAI Grok, DeepSeek, Qwen, NVIDIA NIM, Groq, OpenRouter, Ollama, or your own Custom API endpoint like OmniRoute or vLLM). For voice transcription, choose between Cloud STT (Deepgram, OpenAI Whisper, Groq, NVIDIA NIM) or 100% Local Offline STT (Parakeet streaming and batch models).
Right Alt in any text field, IDE, browser, or terminal."comma", "new line"), bulleted lists, and personal vocabulary casing without hallucinations."um", "uh"), formats code snippets, and matches the tone of your current application.The same gesture (double-tapping Right Alt) stops recording and triggers typing.
Wispr Flow is a proprietary, subscription-based commercial application primarily focused on macOS and Windows. AutoType is purpose-built for Linux users who want complete control over their desktop, privacy, and models.
| Feature | AutoType | Wispr Flow |
|---|---|---|
| Linux Native | Built for Linux Mint XFCE & X11 desktops | Not a primary platform |
| Pricing | 100% Free & Open Source. Pay only your own API provider rates (fractions of a cent) or run completely free offline | Monthly/annual paid subscription |
| Offline Speech | Yes — 100% offline Local STT using Parakeet GGUF models | Cloud only |
| LLM Cleaning Provider | Bring your own: OpenAI, Anthropic, Grok, DeepSeek, Qwen, NVIDIA NIM, Groq, Ollama, or Custom endpoints | Proprietary locked model |
| Custom Prompts & Vocab | Full control over system prompts, vocabulary lists, and style modes | Limited custom dictionary |
| Clipboard Safety | Automatically backs up and restores previous clipboard contents | Overwrites clipboard |
| Privacy | Choose 100% on-device processing (Local STT + Local LLM) with zero data leaving your PC | Audio streamed to third-party cloud |
[ Microphone ]
│
▼
[ Speech-to-Text Layer ]
┌──────────────────┴──────────────────┐
▼ ▼
Cloud STT Local STT
• Deepgram Flux / Nova-3 • Parakeet Stream (120M EOU)
• OpenAI Whisper • Parakeet Batch (0.6B GGUF)
• Groq Whisper (<300ms) • Zero internet required
• NVIDIA NIM Cloud ASR
• Custom OpenAI-compatible ASR
└──────────────────┬──────────────────┘
│ (Raw Transcript)
▼
[ Deterministic Normalizer ]
• Spoken punctuation ("period", "question mark", "open quote")
• Formatting ("new line", "bullet point", "numbered list")
• Vocabulary casing ("Power BI", "FastAPI", "DataFrame")
│
▼
[ LLM Text Cleaning Layer ]
• OpenAI (ChatGPT) • NVIDIA NIM
• Anthropic (Claude) • Groq (Llama-3.3)
• xAI (Grok) • OpenRouter
• DeepSeek • Ollama (Local)
• Qwen (Alibaba DashScope) • Custom API (OmniRoute, vLLM)
(Skipped entirely when using Voice Command or Mode: "raw")
│
▼
[ Desktop Output Layer ]
• Window detection (xdotool) → adaptive style (code / email / chat)
• Auto-paste into focused window (Ctrl+V / Ctrl+Shift+V)
• Clipboard restored to previous content
AutoType eliminates hardcoded vendor lock-in. You can use any major AI provider or bring your own self-hosted server:
| Provider | Protocol | Default Model | Typical Use Case |
|---|---|---|---|
| OpenAI (ChatGPT) | /v1/chat/completions | gpt-4o-mini | Reliable, fast, high-quality standard |
| Anthropic (Claude) | /v1/messages (Native) | claude-3-5-haiku-20241022 | Exceptional instruction following & nuance |
| xAI (Grok) | /v1/chat/completions | grok-2-mini | High-speed processing with modern reasoning |
| DeepSeek | /v1/chat/completions | deepseek-chat | State-of-the-art cost efficiency |
| Qwen (Alibaba Cloud) | /v1/chat/completions | qwen-plus | Strong multilingual and reasoning capability |
| NVIDIA NIM | /v1/chat/completions | meta/llama-3.3-70b-instruct | Enterprise-grade accelerated inference |
| Groq | /v1/chat/completions | llama-3.3-70b-versatile | Ultra-low latency (~200ms turnaround) |
| OpenRouter | /v1/chat/completions | meta-llama/llama-3.3-70b-instruct | Unified gateway to hundreds of open/closed models |
| Ollama (Local) | /v1/chat/completions | llama3.2 | 100% private, on-device local cleaning |
| Custom API Provider | /v1/chat/completions | Custom / auto | Bring your own endpoint (e.g. OmniRoute, vLLM, LiteLLM) |
Configure your provider easily via the graphical Settings GUI or by editing .env.
AutoType cleanly separates speech recognition into two distinct pipelines:
For maximum vocabulary coverage and rapid setup using cloud APIs:
flux-general-en) and batch accuracy with Nova-3 (nova-3).whisper-1).nvidia/parakeet-ctc-1.1b-asr)./v1/audio/transcriptions endpoint.For air-gapped environments or total data sovereignty:
parakeet-cli and an End-of-Utterance (EOU) 120M GGUF model. Audio is transcribed as you speak.transcribe-cli and a 0.6B GGUF model. Transcribes immediately upon releasing the hotkey.sudo apt update
sudo apt install \
python3 python3-venv python3-pip \
python3-gi python3-gi-cairo gir1.2-gtk-3.0 gir1.2-handy-1 \
libportaudio2 \
xclip xdotool libnotify-bin pulseaudio-utils
git clone https://github.com/premkumar-1122/AutoType.git
cd AutoType
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
.env)cp .env.example .env
chmod 600 .env
Open .env in your editor or configure everything via the GTK Settings GUI.
# Cleaning Provider
LLM_PROVIDER=openai
LLM_API_KEY=sk-...
LLM_BASE_URL=https://api.openai.com/v1
LLM_MODEL=gpt-4o-mini
# Speech Engine
STT_BACKEND=deepgram
DEEPGRAM_API_KEY=your_deepgram_key
# Cleaning Provider
LLM_PROVIDER=anthropic
LLM_API_KEY=sk-ant-...
LLM_BASE_URL=https://api.anthropic.com/v1
LLM_MODEL=claude-3-5-haiku-20241022
# Speech Engine
STT_BACKEND=groq
CLOUD_STT_API_KEY=gsk_...
CLOUD_STT_BASE_URL=https://api.groq.com/openai/v1
CLOUD_STT_MODEL=whisper-large-v3-turbo
# Cleaning Provider
LLM_PROVIDER=ollama
LLM_BASE_URL=http://localhost:11434/v1
LLM_MODEL=llama3.2
# Speech Engine (100% Offline)
STT_BACKEND=parakeet_stream
PARAKEET_STREAM_BINARY=/path/to/parakeet-cli
PARAKEET_STREAM_MODEL=/path/to/parakeet-eou-120m-q8_0.gguf
# Custom API Endpoint
LLM_PROVIDER=custom
LLM_BASE_URL=http://127.0.0.1:20128/v1
LLM_API_KEY=your_key
LLM_MODEL=auto
STT_BACKEND=deepgram
DEEPGRAM_API_KEY=your_deepgram_key
(Note: Legacy OMNIROUTE_* environment variables are automatically mapped to LLM_* for backwards compatibility).
AutoType comes with an intuitive, native GTK3/libhandy settings manager:
./open-settings.sh
(Or right-click the tray icon and select Settings…).
clean, raw, smart, code), configure personal vocabulary, and restart the daemon.Transcribing... → Cleaning... → Done, and types the result directly into your active window.Start your dictation with your configured COMMAND_PREFIX (default: "computer") to trigger instant actions:
| Voice Command | Action |
|---|---|
computer cancel | Discards the current recording (also accepts "never mind", "scratch that") |
computer raw ... | One-shot raw paste (bypasses LLM rewriting) |
computer clean ... | Standard AI cleanup |
computer smart ... | Automatically matches the style of the focused application |
computer professional ... | Strict business/formal tone |
computer casual ... | Friendly, informal messaging tone |
computer email ... | Structured email layout |
computer code ... | Preserves camelCase, snake_case, paths, and syntax literally |
Normalizer runs deterministically on your device:
"comma", "period", "question mark", "exclamation mark", "colon", "semicolon", "dash", "open quote" / "close quote", "new line", "new paragraph"."bullet point" or "numbered list ... first ... second ..." to generate formatted lists.source .venv/bin/activate
python app.py
nohup .venv/bin/python app.py >/dev/null 2>&1 &
Open Settings → Session and Startup → Application Autostart.
Click Add.
Set Command to:
/full/path/to/AutoType/.venv/bin/python /full/path/to/AutoType/app.py
default Capture vs PipeWire: On many Linux Mint installations, targeting ALSA's default device results in zero audio capture (digital silence). AutoType explicitly uses PipeWire ALSA routing when MIC_DEVICE is left empty, ensuring seamless dynamic switching between internal microphones and external headsets.AutoType/
├── app.py # Daemon lifecycle, double-tap hotkey listener, pipeline orchestrator
├── config.py # Central configuration loader (.env parser with fallback logic)
├── settings_gui.py # Native GTK3/libhandy graphical configuration app
├── open-settings.sh # Convenience launcher for Settings GUI
│
├── audio/ # Microphone stream capture, level metering, Bluetooth routing
├── stt/ # Speech-to-Text backends
│ ├── deepgram.py # Cloud STT: Deepgram Nova-3 batch
│ ├── flux.py # Cloud STT: Deepgram Flux real-time streaming
│ ├── cloud_whisper.py # Cloud STT: OpenAI Whisper / Groq / NVIDIA NIM / Custom endpoints
│ ├── parakeet.py # Local STT: Offline batch (transcribe-cli 0.6B)
│ ├── parakeet_stream.py # Local STT: Offline real-time streaming (parakeet-cli 120M)
│ └── factory.py # Dynamic STT transcriber dispatch
│
├── llm/ # Provider-Agnostic LLM Cleaning
│ ├── processor.py # Multi-provider client (OpenAI, Anthropic, Grok, DeepSeek, Qwen, etc.)
│ ├── prompts.py # Style definitions (smart/clean/code/email/casual) and system prompts
│ ├── normalizer.py # Deterministic text normalizer (punctuation, lists, spacing)
│ ├── vocabulary.py # Personal vocabulary capitalization matcher
│ └── omniroute.py # Backwards-compatible alias for LLMProcessor
│
├── context/ # Window detection, voice commands, clipboard history
├── desktop/ # X11 clipboard integration, key simulation, window title parsing
├── ui/ # Floating taskbar overlay pill, tray icon, profile storage
└── data/ # Local history logs, custom prompt, personal vocabulary
.env and data/gui_settings.json store your keys locally with restricted permissions (chmod 600) and are strictly ignored by git.parakeet_stream or parakeet) with Local LLM Cleaning (ollama) guarantees zero audio or text data ever leaves your computer.This project is licensed under the MIT License.
Voice typing for Linux Mint XFCE. Double-tap Right Alt, speak, and clean formatted text is instantly pasted into whatever window you are using.
AutoType is a fast, flexible, open-source Linux alternative to Wispr Flow. It runs as a lightweight, silent background daemon: listening for a single hotkey gesture, converting speech to text, intelligently cleaning the transcription, and typing it into your active app.
Bring your own providers: AutoType is completely provider-agnostic. Choose any popular AI provider for text cleaning (OpenAI, Anthropic Claude, xAI Grok, DeepSeek, Qwen, NVIDIA NIM, Groq, OpenRouter, Ollama, or your own Custom API endpoint like OmniRoute or vLLM). For voice transcription, choose between Cloud STT (Deepgram, OpenAI Whisper, Groq, NVIDIA NIM) or 100% Local Offline STT (Parakeet streaming and batch models).
Right Alt in any text field, IDE, browser, or terminal."comma", "new line"), bulleted lists, and personal vocabulary casing without hallucinations."um", "uh"), formats code snippets, and matches the tone of your current application.The same gesture (double-tapping Right Alt) stops recording and triggers typing.
Wispr Flow is a proprietary, subscription-based commercial application primarily focused on macOS and Windows. AutoType is purpose-built for Linux users who want complete control over their desktop, privacy, and models.
| Feature | AutoType | Wispr Flow |
|---|---|---|
| Linux Native | Built for Linux Mint XFCE & X11 desktops | Not a primary platform |
| Pricing | 100% Free & Open Source. Pay only your own API provider rates (fractions of a cent) or run completely free offline | Monthly/annual paid subscription |
| Offline Speech | Yes — 100% offline Local STT using Parakeet GGUF models | Cloud only |
| LLM Cleaning Provider | Bring your own: OpenAI, Anthropic, Grok, DeepSeek, Qwen, NVIDIA NIM, Groq, Ollama, or Custom endpoints | Proprietary locked model |
| Custom Prompts & Vocab | Full control over system prompts, vocabulary lists, and style modes | Limited custom dictionary |
| Clipboard Safety | Automatically backs up and restores previous clipboard contents | Overwrites clipboard |
| Privacy | Choose 100% on-device processing (Local STT + Local LLM) with zero data leaving your PC | Audio streamed to third-party cloud |
[ Microphone ]
│
▼
[ Speech-to-Text Layer ]
┌──────────────────┴──────────────────┐
▼ ▼
Cloud STT Local STT
• Deepgram Flux / Nova-3 • Parakeet Stream (120M EOU)
• OpenAI Whisper • Parakeet Batch (0.6B GGUF)
• Groq Whisper (<300ms) • Zero internet required
• NVIDIA NIM Cloud ASR
• Custom OpenAI-compatible ASR
└──────────────────┬──────────────────┘
│ (Raw Transcript)
▼
[ Deterministic Normalizer ]
• Spoken punctuation ("period", "question mark", "open quote")
• Formatting ("new line", "bullet point", "numbered list")
• Vocabulary casing ("Power BI", "FastAPI", "DataFrame")
│
▼
[ LLM Text Cleaning Layer ]
• OpenAI (ChatGPT) • NVIDIA NIM
• Anthropic (Claude) • Groq (Llama-3.3)
• xAI (Grok) • OpenRouter
• DeepSeek • Ollama (Local)
• Qwen (Alibaba DashScope) • Custom API (OmniRoute, vLLM)
(Skipped entirely when using Voice Command or Mode: "raw")
│
▼
[ Desktop Output Layer ]
• Window detection (xdotool) → adaptive style (code / email / chat)
• Auto-paste into focused window (Ctrl+V / Ctrl+Shift+V)
• Clipboard restored to previous content
AutoType eliminates hardcoded vendor lock-in. You can use any major AI provider or bring your own self-hosted server:
| Provider | Protocol | Default Model | Typical Use Case |
|---|---|---|---|
| OpenAI (ChatGPT) | /v1/chat/completions | gpt-4o-mini | Reliable, fast, high-quality standard |
| Anthropic (Claude) | /v1/messages (Native) | claude-3-5-haiku-20241022 | Exceptional instruction following & nuance |
| xAI (Grok) | /v1/chat/completions | grok-2-mini | High-speed processing with modern reasoning |
| DeepSeek | /v1/chat/completions | deepseek-chat | State-of-the-art cost efficiency |
| Qwen (Alibaba Cloud) | /v1/chat/completions | qwen-plus | Strong multilingual and reasoning capability |
| NVIDIA NIM | /v1/chat/completions | meta/llama-3.3-70b-instruct | Enterprise-grade accelerated inference |
| Groq | /v1/chat/completions | llama-3.3-70b-versatile | Ultra-low latency (~200ms turnaround) |
| OpenRouter | /v1/chat/completions | meta-llama/llama-3.3-70b-instruct | Unified gateway to hundreds of open/closed models |
| Ollama (Local) | /v1/chat/completions | llama3.2 | 100% private, on-device local cleaning |
| Custom API Provider | /v1/chat/completions | Custom / auto | Bring your own endpoint (e.g. OmniRoute, vLLM, LiteLLM) |
Configure your provider easily via the graphical Settings GUI or by editing .env.
AutoType cleanly separates speech recognition into two distinct pipelines:
For maximum vocabulary coverage and rapid setup using cloud APIs:
flux-general-en) and batch accuracy with Nova-3 (nova-3).whisper-1).nvidia/parakeet-ctc-1.1b-asr)./v1/audio/transcriptions endpoint.For air-gapped environments or total data sovereignty:
parakeet-cli and an End-of-Utterance (EOU) 120M GGUF model. Audio is transcribed as you speak.transcribe-cli and a 0.6B GGUF model. Transcribes immediately upon releasing the hotkey.sudo apt update
sudo apt install \
python3 python3-venv python3-pip \
python3-gi python3-gi-cairo gir1.2-gtk-3.0 gir1.2-handy-1 \
libportaudio2 \
xclip xdotool libnotify-bin pulseaudio-utils
git clone https://github.com/premkumar-1122/AutoType.git
cd AutoType
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
.env)cp .env.example .env
chmod 600 .env
Open .env in your editor or configure everything via the GTK Settings GUI.
# Cleaning Provider
LLM_PROVIDER=openai
LLM_API_KEY=sk-...
LLM_BASE_URL=https://api.openai.com/v1
LLM_MODEL=gpt-4o-mini
# Speech Engine
STT_BACKEND=deepgram
DEEPGRAM_API_KEY=your_deepgram_key
# Cleaning Provider
LLM_PROVIDER=anthropic
LLM_API_KEY=sk-ant-...
LLM_BASE_URL=https://api.anthropic.com/v1
LLM_MODEL=claude-3-5-haiku-20241022
# Speech Engine
STT_BACKEND=groq
CLOUD_STT_API_KEY=gsk_...
CLOUD_STT_BASE_URL=https://api.groq.com/openai/v1
CLOUD_STT_MODEL=whisper-large-v3-turbo
# Cleaning Provider
LLM_PROVIDER=ollama
LLM_BASE_URL=http://localhost:11434/v1
LLM_MODEL=llama3.2
# Speech Engine (100% Offline)
STT_BACKEND=parakeet_stream
PARAKEET_STREAM_BINARY=/path/to/parakeet-cli
PARAKEET_STREAM_MODEL=/path/to/parakeet-eou-120m-q8_0.gguf
# Custom API Endpoint
LLM_PROVIDER=custom
LLM_BASE_URL=http://127.0.0.1:20128/v1
LLM_API_KEY=your_key
LLM_MODEL=auto
STT_BACKEND=deepgram
DEEPGRAM_API_KEY=your_deepgram_key
(Note: Legacy OMNIROUTE_* environment variables are automatically mapped to LLM_* for backwards compatibility).
AutoType comes with an intuitive, native GTK3/libhandy settings manager:
./open-settings.sh
(Or right-click the tray icon and select Settings…).
clean, raw, smart, code), configure personal vocabulary, and restart the daemon.Transcribing... → Cleaning... → Done, and types the result directly into your active window.Start your dictation with your configured COMMAND_PREFIX (default: "computer") to trigger instant actions:
| Voice Command | Action |
|---|---|
computer cancel | Discards the current recording (also accepts "never mind", "scratch that") |
computer raw ... | One-shot raw paste (bypasses LLM rewriting) |
computer clean ... | Standard AI cleanup |
computer smart ... | Automatically matches the style of the focused application |
computer professional ... | Strict business/formal tone |
computer casual ... | Friendly, informal messaging tone |
computer email ... | Structured email layout |
computer code ... | Preserves camelCase, snake_case, paths, and syntax literally |
Normalizer runs deterministically on your device:
"comma", "period", "question mark", "exclamation mark", "colon", "semicolon", "dash", "open quote" / "close quote", "new line", "new paragraph"."bullet point" or "numbered list ... first ... second ..." to generate formatted lists.source .venv/bin/activate
python app.py
nohup .venv/bin/python app.py >/dev/null 2>&1 &
Open Settings → Session and Startup → Application Autostart.
Click Add.
Set Command to:
/full/path/to/AutoType/.venv/bin/python /full/path/to/AutoType/app.py
default Capture vs PipeWire: On many Linux Mint installations, targeting ALSA's default device results in zero audio capture (digital silence). AutoType explicitly uses PipeWire ALSA routing when MIC_DEVICE is left empty, ensuring seamless dynamic switching between internal microphones and external headsets.AutoType/
├── app.py # Daemon lifecycle, double-tap hotkey listener, pipeline orchestrator
├── config.py # Central configuration loader (.env parser with fallback logic)
├── settings_gui.py # Native GTK3/libhandy graphical configuration app
├── open-settings.sh # Convenience launcher for Settings GUI
│
├── audio/ # Microphone stream capture, level metering, Bluetooth routing
├── stt/ # Speech-to-Text backends
│ ├── deepgram.py # Cloud STT: Deepgram Nova-3 batch
│ ├── flux.py # Cloud STT: Deepgram Flux real-time streaming
│ ├── cloud_whisper.py # Cloud STT: OpenAI Whisper / Groq / NVIDIA NIM / Custom endpoints
│ ├── parakeet.py # Local STT: Offline batch (transcribe-cli 0.6B)
│ ├── parakeet_stream.py # Local STT: Offline real-time streaming (parakeet-cli 120M)
│ └── factory.py # Dynamic STT transcriber dispatch
│
├── llm/ # Provider-Agnostic LLM Cleaning
│ ├── processor.py # Multi-provider client (OpenAI, Anthropic, Grok, DeepSeek, Qwen, etc.)
│ ├── prompts.py # Style definitions (smart/clean/code/email/casual) and system prompts
│ ├── normalizer.py # Deterministic text normalizer (punctuation, lists, spacing)
│ ├── vocabulary.py # Personal vocabulary capitalization matcher
│ └── omniroute.py # Backwards-compatible alias for LLMProcessor
│
├── context/ # Window detection, voice commands, clipboard history
├── desktop/ # X11 clipboard integration, key simulation, window title parsing
├── ui/ # Floating taskbar overlay pill, tray icon, profile storage
└── data/ # Local history logs, custom prompt, personal vocabulary
.env and data/gui_settings.json store your keys locally with restricted permissions (chmod 600) and are strictly ignored by git.parakeet_stream or parakeet) with Local LLM Cleaning (ollama) guarantees zero audio or text data ever leaves your computer.This project is licensed under the MIT License.