DictionLabs/Diction

Open Source self-hosted alternative to WisprFlow - iOS voice AI keyboard

Go

216

180 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Diction - self-hosted voice and AI keyboard for iPhone (r/coolgithubprojects)

iOS keyboard with dictation, voice editing and a full QWERTY. The speech models and the LLM run on your own server, or on the phone. Both are free, no limits. https://github.com/DictionLabs/Diction https://apps.apple.com/app/id6759807364

1

Oct 6, 2026

README

Diction

The iPhone keyboard
with voice and AI you can self-host.


A full QWERTY that learns your words.
On-device, cloud, or your own server.

Type 5x faster with your voice Switch keyboard and record Hold for edit mode

Download on the App Store

Website • Self-Hosting Guide • Privacy Policy

App Store Latest gateway release License: MIT

Gateway tests Coverage Docker pulls

Gateway image size Hugging Face

Contributors


What is Diction?

Diction is an iOS keyboard that transcribes speech to text directly in any app. Tap the mic, speak, text lands in the field. No switching apps, no copy-paste.

This repo is the open-source gateway: a Go service that sits between the iOS keyboard and your speech-to-text backend. It handles the WebSocket streaming protocol, AES-256-GCM end-to-end encryption, and optional LLM cleanup. The iOS app is on the App Store; the gateway is what you self-host.

  • Self-hosted in one command. docker compose up and paste the URL into the app. Your server, your models, your data.
  • Model-agnostic. The gateway speaks the OpenAI transcription API spec. Point it at any speech-to-text backend that implements it. Your model, your stack.
  • On-device option. On-device models run locally on the iPhone. No gateway needed for that mode.
  • Encrypted in transit. AES-256-GCM with X25519 key exchange. Same primitives used by Signal and WireGuard.
  • Zero tracking. No analytics, no telemetry, no data collection. Audit the source yourself.
  • Free and unlimited. Self-hosted and on-device modes have no caps, no rate limits, no expiry.

Self-Hosting

Diction speaks the OpenAI transcription API (POST /v1/audio/transcriptions) directly, so any compatible Whisper server works without the gateway. The gateway adds a WebSocket layer for live streaming — audio is transcribed as you speak, so by the time you tap stop the result is already back. For longer dictations the difference is noticeable; for short phrases it barely matters. The gateway also handles end-to-end encryption and optional LLM cleanup.

Full walkthrough with screenshots: How to Set Up Diction - the self-hosted speech-to-text alternative to Wispr Flow

Requirements:

  • An NVIDIA GPU gets you the setup below. No GPU? Skip to No GPU? for the CPU path.
  • Any machine that can run Docker: Linux box, NUC, home server, VPS.
  • iPhone running iOS 17.0 or later.

Step 1 - Write the Compose File

Install the NVIDIA Container Toolkit on the host first. Create a folder for the stack and save this as docker-compose.yml:

services:
  parakeet:
    image: dictionlabs/parakeet:latest-int8
    container_name: diction-parakeet
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    healthcheck:
      test: ["CMD", "bash", "-c", "echo > /dev/tcp/localhost/5092"]
      interval: 30s
      timeout: 5s
      retries: 3
      start_period: 30s

  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    depends_on:
      - parakeet
    environment:
      DEFAULT_MODEL: parakeet-v3

ghcr.io/omachala/diction-gateway is retired as of v13.0 (2026-09-11). Images already published there stay pullable and are not going anywhere, but no new versions will land on it - that ref is frozen at v12.0. If you are pulling it, switch to dictionlabs/gateway to keep receiving updates:

-    image: ghcr.io/omachala/diction-gateway:latest
+    image: dictionlabs/gateway:latest

Images are published under Diction Labs on Docker Hub (dictionlabs/gateway, canonical) and GHCR (ghcr.io/dictionlabs/gateway).

Model weights are baked into the parakeet image, so there's nothing to download on first start. DEFAULT_MODEL: parakeet-v3 covers 25 European languages - see Swap the Speech Model for Whisper models and other languages.

Step 2 - Start the Stack

docker compose up -d
docker compose logs -f          # watch progress
docker compose ps               # check status

Expected:

NAME                     STATUS
diction-gateway          Up 30 seconds
diction-parakeet         Up 30 seconds (healthy)
ErrorFix
pull access denied on gateway imagedocker logout and retry - a stale login to Docker Hub or ghcr.io can shadow the anonymous pull
could not select device driver "nvidia"NVIDIA Container Toolkit isn't installed, or Docker wasn't restarted after installing it
Gateway exits immediatelyParakeet container failed - check its logs

Step 3 - Test the Server

Generate a test audio file (macOS):

say -o test.aiff "Hello from my home server"

Or record a voice memo on your phone and AirDrop it over.

curl -X POST http://localhost:8080/v1/audio/transcriptions \
  -F "file=@test.aiff" \
  -F "model=parakeet-v3"
{"text":"Hello from my home server."}
# Check timing headers
curl -sS -D - -o /dev/null \
  -X POST http://localhost:8080/v1/audio/transcriptions \
  -F "file=@test.aiff" -F "model=parakeet-v3" | grep -i diction

That returns three headers, verified against a live gateway:

HeaderMeaning
X-Diction-Whisper-Msspeech model inference latency in milliseconds
X-Diction-Route-Modelwhich backend actually served the request
X-Diction-Route-Langdetected language, empty when detection didn't run

X-Diction-LLM-Ms is added when AI cleanup runs. X-Diction-Route-Model is the quickest way to confirm a model switch took effect, since an unrecognised model name silently falls back to DEFAULT_MODEL.

ResponseCause
Connection refusedGateway not running - docker compose ps
502 Bad GatewayParakeet unreachable, still loading, or DEFAULT_MODEL names a service that isn't running
400 Bad RequestUnsupported response_format (only json and text)
404 Not FoundURL typo - path must be exactly /v1/audio/transcriptions
OOM / container crashNot enough VRAM - Parakeet INT8 needs about 2 GB

Step 4 - Find Your Server's IP

macOS:

ipconfig getifaddr en0
# or
ifconfig | grep 'inet ' | grep -v 127.0.0.1

Linux:

hostname -I | awk '{print $1}'

Windows:

ipconfig | findstr IPv4

Pick the 192.168.x.x or 10.x.x.x address. Ignore anything starting with 100. - that's Tailscale.

Set a DHCP reservation in your router so the IP doesn't change on reboot. Or use Tailscale for a stable address that follows the machine anywhere.

Step 5 - Connect the App

Install Diction on your iPhone. On first launch:

  1. Settings → General → Keyboard → Keyboards → Add New Keyboard → Diction
  2. Tap Diction in the list → enable Allow Full Access
  3. Grant microphone access when prompted

Point it at your server:

  1. Open Diction → Preferences → Mode → Self-Hosted
  2. Enter your endpoint: http://192.168.1.42:8080 (your IP from Step 4)
  3. Tap Test connection - you should get a green check within a second

To dictate: open any app, tap a text field, long-press the globe icon (bottom-left of the iOS keyboard), pick Diction, tap the mic, speak, release.

Reach From Anywhere

Tailscale (recommended)

Tailscale creates a private WireGuard mesh between your devices. Install it on the server and iPhone, sign in to the same account, and use the 100.x.x.x Tailscale IP as your Diction endpoint. Works on cellular, café WiFi, anywhere. Free for personal use.

Cloudflare Tunnel (public URL, no port forwarding)

Add to your compose file:

  cloudflared:
    image: cloudflare/cloudflared:latest
    container_name: diction-cloudflared
    restart: unless-stopped
    command: tunnel --no-autoupdate run
    environment:
      TUNNEL_TOKEN: "${CLOUDFLARE_TUNNEL_TOKEN}"

Create a tunnel in the Cloudflare Zero Trust dashboard, grab the token, add it to .env, route the public hostname to http://gateway:8080. Free tier. Note: transcripts pass through Cloudflare's network (HTTPS-encrypted, but a third party is in the path).

ngrok (quick testing)

ngrok http 8080

Free tier URLs change on restart - good for a demo, not daily use.


No GPU? (CPU-only Whisper)

No NVIDIA GPU on the box you're using? Run Whisper on CPU instead. Slower than Parakeet, but it runs anywhere Docker does.

services:
  # The whisper server runs as uid 1000. A volume left behind by the older root-based
  # images is owned by root, so without this the server cannot write and every model
  # download fails with a permission error. No-op on a fresh install.
  whisper-models-init:
    image: dictionlabs/whisper-server:latest-cpu
    user: root
    volumes:
      - whisper-models:/cache
    entrypoint: ["sh", "-c", "chown -R 1000:1000 /cache"]
    restart: "no"

  whisper-small:
    image: dictionlabs/whisper-server:latest-cpu
    container_name: diction-whisper-small
    restart: unless-stopped
    depends_on:
      whisper-models-init:
        condition: service_completed_successfully
    volumes:
      - whisper-models:/home/ubuntu/.cache/huggingface/hub
    healthcheck:
      test: ["CMD", "curl", "-fsS", "-o", "/dev/null", "http://localhost:8000/health"]
      interval: 30s
      timeout: 5s
      retries: 3
      start_period: 120s

  # The whisper server never downloads a model on demand: it answers "not installed
  # locally" instead of fetching. Without this one-shot pull the stack starts cleanly
  # and then fails every transcription. Idempotent, exits once the model is on disk.
  whisper-small-pull:
    image: dictionlabs/whisper-server:latest-cpu
    depends_on:
      whisper-small:
        condition: service_healthy
    entrypoint: ["curl", "-fsS", "-X", "POST",
                 "http://whisper-small:8000/v1/models/DictionLabs/whisper-small-ct2"]
    restart: "no"

  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    depends_on:
      - whisper-small
    environment:
      DEFAULT_MODEL: small

volumes:
  whisper-models:

The whisper-models volume persists the model weights (~500 MB for small) so they survive container rebuilds. The first up takes a few minutes longer than later ones while whisper-small-pull downloads the model - it exits as soon as that finishes, and subsequent starts skip straight past it.

docker compose up -d
ErrorFix
exec format error on Apple SiliconEnable Rosetta in Docker Desktop → Settings → General
health: starting for > 3 minutesModel still downloading - docker compose logs -f whisper-small
Gateway exits immediatelyWhisper container failed - check its logs

Test it the same way as Step 3, with -F "model=small" instead of parakeet-v3. See Swap the Speech Model below for medium and large-v3-turbo.


Swap the Speech Model

Pick a compose profile and set DEFAULT_MODEL on the gateway to match:

DEFAULT_MODELCompose profileService nameWeights servedRAMNotes
smallsmallwhisper-smallDictionLabs/whisper-small-ct2~850 MBBest for CPU
mediummediumwhisper-mediumDictionLabs/whisper-medium-ct2~2.1 GBMore accurate, slower on CPU
large-v3-turbolargewhisper-large-turboDictionLabs/whisper-large-v3-turbo-ct2~2.3 GBHighest accuracy; slow on CPU, fast on GPU
parakeet-v3parakeetparakeetbaked into the image~2 GB25 European languages; fast on CPU, faster on GPU
docker compose --profile small up -d

DEFAULT_MODEL and the service name must both match the table: the gateway resolves backends by Docker hostname, so if the named service isn't running every request fails with 502.

An unrecognised model name is not an error. The gateway falls back to DEFAULT_MODEL, so a typo transcribes successfully with the wrong model rather than telling you. Check the X-Diction-Route-Model response header to see which backend actually served a request.

The weights column is informational. You do not set it anywhere: the gateway names the model when it forwards each request. The whisper server does not fetch models on demand, so the compose file ships a one-shot pull service per profile that installs the model before first use. Those are our own CTranslate2 builds of OpenAI's checkpoints, published at huggingface.co/DictionLabs, so the models your server pulls come from a namespace we control rather than a third party's conversion.

docker compose up -d   # recreates only the changed container

NVIDIA GPU: More Languages (large-v3-turbo)

Already set up Parakeet from Step 1? That covers 25 European languages. For the other 74, swap in Whisper large-v3-turbo running on GPU instead:

Parakeet TDT 0.6B v3Whisper Large-v3-turbo
WER (English)~6.3%7.4%
Latency (GPU)Sub-secondUnder 2s
VRAM~2 GB~2.3 GB
Languages25 European99
services:
  whisper-large-turbo:
    image: dictionlabs/whisper-server:latest-cuda
    container_name: diction-whisper-large-turbo
    restart: unless-stopped
    volumes:
      - whisper-models:/home/ubuntu/.cache/huggingface/hub
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    depends_on:
      - whisper-large-turbo
    environment:
      DEFAULT_MODEL: large-v3-turbo

volumes:
  whisper-models:

First boot downloads ~1.6 GB of model weights into the volume. Subsequent starts are instant.


Already Have a Voice Server?

Keep it. Use CUSTOM_BACKEND_URL to put the Diction Gateway in front of your existing server for WebSocket streaming and end-to-end encryption:

services:
  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    environment:
      CUSTOM_BACKEND_URL: http://your-existing-server:8000
      CUSTOM_BACKEND_MODEL: DictionLabs/whisper-small-ct2
VariableDescription
CUSTOM_BACKEND_AUTHAuthorization header forwarded to your backend, e.g. Bearer sk-xxx
CUSTOM_BACKEND_NEEDS_WAVSet to "true" if your backend only accepts WAV - the gateway converts with ffmpeg
CUSTOM_BACKEND_CANONICAL_IDHuggingFace-style ID advertised via /v1/models (default: CUSTOM_BACKEND_MODEL, or the literal custom if that is unset too)

AI Cleanup and Voice Editing (BYO LLM)

The gateway passes transcripts through any OpenAI-compatible LLM before returning them. You say "so um basically the meeting went well and uh they agreed to the timeline." The LLM returns "The meeting went well. They agreed to the timeline."

Enable the AI Companion toggle in the app. The gateway forwards the transcript to {LLM_BASE_URL}/chat/completions with your prompt, then returns the cleaned text. If the LLM fails, the raw transcript is returned -- dictation never breaks.

When LLM is configured, two additional text routes become available (requires TEXT_ROUTES_OPEN=true):

  • POST /v1/text/process?intent=edit -- apply a spoken voice instruction to text
  • POST /v1/text/process?intent=edit-selected -- apply an instruction to a selected text range
  • POST /v1/text/suggest -- return 2-3 alternative phrasings (always soft-fails)
  • POST /v1/text/summarize -- one-line summary of a saved note (always soft-fails)

Transcript cleanup also honours a "formatting" flag in the request context: when set (or omitted, which means on), LLM_PROMPT_FORMATTING is appended to the cleanup prompt so spoken lists and topic changes come back as line breaks and plain-text bullets.

These routes enable AI parity with Diction One on self-hosted gateways: voice-edit and suggestions work the same way, though results depend on the quality of your chosen model. Smaller models (under 7B) often do not follow edit instructions reliably -- 7B or larger is recommended for editing.

See AGENTS.md for the full wire format.

VariableRequiredDescription
LLM_BASE_URLYesOpenAI-compatible endpoint, e.g. https://api.openai.com/v1
LLM_MODELYesModel identifier, e.g. gpt-4o-mini
LLM_API_KEYNoBearer token. Not needed for local Ollama.
LLM_PROMPTNoSystem prompt for transcript cleanup. String or a file path starting with / (mount via volume). Defaults to the built-in cleanup prompt when empty.
LLM_PROMPT_EDITNoSystem prompt for voice-edit intent (?intent=edit). Defaults to the built-in edit prompt.
LLM_PROMPT_EDIT_SELECTEDNoSystem prompt for edit-selected intent. Defaults to the built-in edit-selected prompt.
LLM_PROMPT_SUGGESTNoSystem prompt for the suggest endpoint. Defaults to the built-in suggest prompt.
LLM_REASONING_EFFORTNoOpenAI-compatible reasoning effort such as none, low, medium, or high. Omitted by default.
LLM_PROMPT_FORMATTINGNoAppended to LLM_PROMPT when the client requests formatting. Defaults to the built-in formatting rules.
LLM_PROMPT_SUMMARYNoSystem prompt for /v1/text/summarize. Defaults to the built-in summary prompt.
TEXT_ROUTES_OPENNoSet to true to open /v1/text/process and /v1/text/suggest when AUTH_ENABLED=false. Default false (routes return 403 until explicitly opened).
DICTION_ENHANCE_TIMEOUT_MSNoHow long the cleanup pass may run while the user is waiting — /v1/audio/transcriptions, and /v1/audio/stream on voice-edit intents. Default 20000 (20s), sized for a local model on CPU. On timeout the raw transcript is returned, so nothing is ever lost. Set 0 for no limit.
DICTION_LIVE_ENHANCE_TIMEOUT_MSNoHow long the cleanup pass may run after the raw text has already been delivered, on /v1/audio/stream?split_enhance=true. Default 8000 (8s). Raising it much past 9s has no effect: the app stops waiting for the enhanced frame at 9s and keeps the raw text. Set 0 for no limit.

Both LLM_BASE_URL and LLM_MODEL must be set or the feature stays off.

Language hint: when the client's context carries a concrete language (not empty, not the auto-detect sentinel "auto"), the cleanup call gets a (Language: xx) hint appended to the user message — the built-in LLM_PROMPT tells the model to write in that language and correct wrong or missing accents/diacritics for it, never to translate. A custom LLM_PROMPT should say what to do with the hint if it matters to you; it is plain text appended after the transcript, not a template variable.

What the cleanup call sends. The transcript comes first, unlabelled, exactly as before. The descriptive context the user configured in the app then follows, one labelled line per item, each omitted entirely when unset — a user with none of them configured sends byte-for-byte the request earlier releases sent:

LineSourceCap
Custom words: a, b (also heard as: c)the user's My Words list50 entries
Tone: ...the user's Tone preset and About You description, mergednone
(Language: xx)the transcript's language, when concrete—

What it deliberately does NOT send, and why you should not add it. The client also supplies the text around the cursor, the earlier transcripts of the session, and the clipboard. The gateway decodes them and forwards none of them to the cleanup prompt. Measured against gpt-oss-20b with the built-in prompt: given a block of three earlier transcripts, the dictation "yeah that sounds good" came back as those three transcripts — the user's actual words gone from their text field. A long clipboard block was appended to the output verbatim. Custom words, tone and profile never did this.

The rule that survived: forward context that describes the speaker, never context that is itself prose the model could emit as the answer. The built-in prompt does say these lines are context and must not appear in the reply; a short prompt does not enforce it. If you write a custom LLM_PROMPT and want to feed it session or clipboard context, you own that trade — test it with a short dictation and a long context block before trusting it.

What the voice-edit call sends. For intent=edit (the user's cursor, nothing selected) the gateway sends Text: <before>‸<after> followed by Instruction: <what the user said>. The ‸ marks the cursor. The model is expected to return the full modified text with the marker removed; the gateway strips any ‸ the model leaves behind, because the app inserts the result verbatim. For intent=edit-selected the selection is the text and the surrounding context follows as Context before: / Context after: lines. An edit request with nothing to edit (no cursor context, or no selection) fails rather than guessing, so the app shows "Couldn't apply edit" instead of typing the user's spoken instruction into their document. A custom LLM_PROMPT_EDIT should account for the marker.

If cleanup keeps returning raw text, check the startup log line — it prints enhance_ms= and live_enhance_ms= — and time your model directly against LLM_BASE_URL. A local model slower than the budget is the usual cause; raise DICTION_ENHANCE_TIMEOUT_MS rather than switching cleanup off.

Behavior change from earlier releases: operators who set LLM_BASE_URL and LLM_MODEL without LLM_PROMPT now receive the built-in cleanup prompt automatically. Previously the gateway logged a warning and sent no system instructions. The default prompt is: "You are a transcript cleanup tool. Fix grammar, punctuation, and remove filler words. If a language is given, write in that language and correct wrong or missing accents or diacritics for it. Never translate. Lines labelled “Custom words” or “Tone” may follow the transcript: they describe the speaker, never part of what you return. Return only the corrected transcript, nothing else."

Option A - Cloud LLM (OpenAI, Groq, etc.)

echo "OPENAI_API_KEY=sk-your-key-here" > .env
  gateway:
    environment:
      DEFAULT_MODEL: small
      LLM_BASE_URL: "https://api.openai.com/v1"
      LLM_API_KEY: "${OPENAI_API_KEY}"
      LLM_MODEL: "gpt-4o-mini"
      LLM_PROMPT: "Clean up this voice transcription. Remove filler words (um, uh, like). Fix punctuation and capitalization. Return only the cleaned text, nothing else."

Docker Compose reads ${OPENAI_API_KEY} from .env automatically. Works with any OpenAI-compatible provider - Groq, Together, Fireworks, Mistral, OpenRouter - swap LLM_BASE_URL and LLM_MODEL.

Option B - Local Ollama (zero cost, fully private)

  ollama:
    image: ollama/ollama:latest
    container_name: diction-ollama
    restart: unless-stopped
    volumes:
      - ollama-models:/root/.ollama

  gateway:
    environment:
      DEFAULT_MODEL: small
      LLM_BASE_URL: "http://ollama:11434/v1"
      LLM_MODEL: "gemma2:9b"
      LLM_PROMPT: "Clean up this voice transcription. Remove filler words. Fix punctuation and capitalization. Return only the cleaned text, nothing else."

volumes:
  whisper-models:
  ollama-models:
docker compose up -d
docker exec diction-ollama ollama pull gemma2:9b
ModelMemoryNotes
gemma2:9b~6 GBBest cleanup quality at this size
qwen2.5:7b~5 GBStrong instruction following
llama3.1:8b~5 GBMost popular, well-tested
gemma3:4b~3 GBFor tighter machines

Models under 7B tend to answer questions about the transcript instead of cleaning it up. 7B or larger recommended.

Testing cleanup

curl -X POST "http://localhost:8080/v1/audio/transcriptions?enhance=true" \
  -F "file=@test.aiff" \
  -F "model=small"
# Confirm LLM fired - look for X-Diction-LLM-Ms in the output
curl -sS -D - -o /dev/null \
  -X POST "http://localhost:8080/v1/audio/transcriptions?enhance=true" \
  -F "file=@test.aiff" -F "model=small" | grep -i diction

Prompt file

Mount a file and point LLM_PROMPT at the path:

  gateway:
    volumes:
      - ./cleanup-prompt.txt:/config/prompt.txt:ro
    environment:
      LLM_PROMPT: "/config/prompt.txt"

If LLM_PROMPT starts with /, the gateway reads it as a file. Otherwise it uses the string directly.


NixOS

The repo ships a flake with a hardened systemd module - no Docker needed.

nix run github:DictionLabs/Diction#diction-gateway

Enable as a service:

{
  inputs.diction.url = "github:DictionLabs/Diction";

  outputs = { nixpkgs, diction, ... }: {
    nixosConfigurations.your-host = nixpkgs.lib.nixosSystem {
      modules = [
        diction.nixosModules.default
        {
          services.diction-gateway = {
            enable = true;
            openFirewall = true;
            # customBackend.url = "http://127.0.0.1:8000";
            # llm.baseUrl = "http://127.0.0.1:11434/v1";
            # llm.model = "gemma2:9b";
            # environmentFile = "/run/secrets/diction-gateway.env";
          };
        }
      ];
    };
  };
}

The unit runs under DynamicUser with ProtectSystem=strict, NoNewPrivileges, and a narrow syscall filter. Use environmentFile for secrets - they don't end up in the world-readable Nix store. Full option list: nix/module.nix.


OpenAI API Compatibility

The gateway implements the OpenAI audio transcription API - any client that works against api.openai.com/v1/audio/transcriptions works against a Diction gateway.

from openai import OpenAI

client = OpenAI(
    base_url="http://your-server:8080/v1",
    api_key="anything",  # not checked when AUTH_ENABLED=false
)

with open("audio.wav", "rb") as f:
    result = client.audio.transcriptions.create(
        file=f,
        model="small",            # or "DictionLabs/whisper-small-ct2"
        response_format="text",
    )
print(result)

Works with the Node SDK, LangChain, Flowise, n8n, or any tool that expects OpenAI's speech API.

Supported:

  • POST /v1/audio/transcriptions - file, model, language, prompt, response_format=json|text
  • GET /v1/models - returns an OpenAI-compatible data[] array plus a providers[] grouping consumed by the iOS app. HuggingFace IDs (DictionLabs/whisper-small-ct2, nvidia/parakeet-tdt-0.6b-v3) and short aliases (small, medium, large-v3-turbo, parakeet-v3) are both accepted. The older Systran/* and deepdml/* ids stay accepted as aliases, so existing scripts keep working.
  • WebSocket /v1/audio/stream - used by the Diction app for low-latency streaming

Not supported:

  • TTS (/v1/audio/speech)
  • response_format=verbose_json|srt|vtt (no word-level timestamps)
  • SSE streaming on REST (use WebSocket /v1/audio/stream instead)
  • Model download/delete (POST/DELETE /v1/models/{id})
  • OpenAI Realtime API (/v1/realtime)

Authentication is off by default (AUTH_ENABLED=false). Pass any non-empty string as the API key from the client - the gateway doesn't check it.

Do not set AUTH_ENABLED=true on a self-hosted deployment. It does not enable a shared secret. The middleware accepts only an Apple App Store JWS validated against the Apple root CA, or an HMAC trial token signed with TRIAL_SECRET. Both exist for Diction One. Turning it on locks you out of your own server. To expose a gateway publicly, put it behind a reverse proxy or tunnel that does the auth, or keep it on a VPN.

Error shape: errors return {"error":"<message>"}, except auth failures, which return {"error":"unauthorized","reason":"...","message":"..."}, not OpenAI's nested {"error":{"message":"...","type":"..."}}. Most SDKs surface these as HTTPError rather than APIError.


Privacy

  • On-device: Everything stays on your phone. No network connection is made.
  • Self-hosted: Audio goes to your server only. Neither the gateway nor the whisper server persists audio - it's transcribed and discarded.
  • AI cleanup enabled: The transcript (plain text, no audio) goes to your configured LLM. If you use Ollama locally, nothing leaves your machine.
  • Diction One (cloud): Audio is transcribed and immediately discarded. Not stored, not used for training.
  • Zero third-party SDKs in the app. No analytics, no tracking, no telemetry.
  • Full Access is required by iOS for any keyboard that makes network requests. What you type on the QWERTY keys never leaves your phone. The app sends your dictation audio, plus the text around your cursor when you edit by voice, and only to the endpoint you configured.

Read the full Privacy Policy.


Diction One

On-device and self-hosted are completely free with no word limits.

If you don't want to run a server, Diction One gives you a fine-tuned cloud model with advanced audio filtering - without the setup. Audio is sent to the Diction endpoint, transcribed, and immediately discarded. Pricing and trial details are in the app.


Contributing

Contributions are welcome. See CONTRIBUTING.md.

License

MIT. See LICENSE.

ios
keyboard
parakeet
speech-to-text
stt
voice
whisper
wispr
wispr-flow
wispr-flow-alternative
wisprflow-alternative

DictionLabs/Diction

Open Source self-hosted alternative to WisprFlow - iOS voice AI keyboard

Go

216

180 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Diction - self-hosted voice and AI keyboard for iPhone (r/coolgithubprojects)

iOS keyboard with dictation, voice editing and a full QWERTY. The speech models and the LLM run on your own server, or on the phone. Both are free, no limits. https://github.com/DictionLabs/Diction https://apps.apple.com/app/id6759807364

1

Oct 6, 2026

README

Diction

The iPhone keyboard
with voice and AI you can self-host.


A full QWERTY that learns your words.
On-device, cloud, or your own server.

Type 5x faster with your voice Switch keyboard and record Hold for edit mode

Download on the App Store

Website • Self-Hosting Guide • Privacy Policy

App Store Latest gateway release License: MIT

Gateway tests Coverage Docker pulls

Gateway image size Hugging Face

Contributors


What is Diction?

Diction is an iOS keyboard that transcribes speech to text directly in any app. Tap the mic, speak, text lands in the field. No switching apps, no copy-paste.

This repo is the open-source gateway: a Go service that sits between the iOS keyboard and your speech-to-text backend. It handles the WebSocket streaming protocol, AES-256-GCM end-to-end encryption, and optional LLM cleanup. The iOS app is on the App Store; the gateway is what you self-host.

  • Self-hosted in one command. docker compose up and paste the URL into the app. Your server, your models, your data.
  • Model-agnostic. The gateway speaks the OpenAI transcription API spec. Point it at any speech-to-text backend that implements it. Your model, your stack.
  • On-device option. On-device models run locally on the iPhone. No gateway needed for that mode.
  • Encrypted in transit. AES-256-GCM with X25519 key exchange. Same primitives used by Signal and WireGuard.
  • Zero tracking. No analytics, no telemetry, no data collection. Audit the source yourself.
  • Free and unlimited. Self-hosted and on-device modes have no caps, no rate limits, no expiry.

Self-Hosting

Diction speaks the OpenAI transcription API (POST /v1/audio/transcriptions) directly, so any compatible Whisper server works without the gateway. The gateway adds a WebSocket layer for live streaming — audio is transcribed as you speak, so by the time you tap stop the result is already back. For longer dictations the difference is noticeable; for short phrases it barely matters. The gateway also handles end-to-end encryption and optional LLM cleanup.

Full walkthrough with screenshots: How to Set Up Diction - the self-hosted speech-to-text alternative to Wispr Flow

Requirements:

  • An NVIDIA GPU gets you the setup below. No GPU? Skip to No GPU? for the CPU path.
  • Any machine that can run Docker: Linux box, NUC, home server, VPS.
  • iPhone running iOS 17.0 or later.

Step 1 - Write the Compose File

Install the NVIDIA Container Toolkit on the host first. Create a folder for the stack and save this as docker-compose.yml:

services:
  parakeet:
    image: dictionlabs/parakeet:latest-int8
    container_name: diction-parakeet
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
    healthcheck:
      test: ["CMD", "bash", "-c", "echo > /dev/tcp/localhost/5092"]
      interval: 30s
      timeout: 5s
      retries: 3
      start_period: 30s

  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    depends_on:
      - parakeet
    environment:
      DEFAULT_MODEL: parakeet-v3

ghcr.io/omachala/diction-gateway is retired as of v13.0 (2026-09-11). Images already published there stay pullable and are not going anywhere, but no new versions will land on it - that ref is frozen at v12.0. If you are pulling it, switch to dictionlabs/gateway to keep receiving updates:

-    image: ghcr.io/omachala/diction-gateway:latest
+    image: dictionlabs/gateway:latest

Images are published under Diction Labs on Docker Hub (dictionlabs/gateway, canonical) and GHCR (ghcr.io/dictionlabs/gateway).

Model weights are baked into the parakeet image, so there's nothing to download on first start. DEFAULT_MODEL: parakeet-v3 covers 25 European languages - see Swap the Speech Model for Whisper models and other languages.

Step 2 - Start the Stack

docker compose up -d
docker compose logs -f          # watch progress
docker compose ps               # check status

Expected:

NAME                     STATUS
diction-gateway          Up 30 seconds
diction-parakeet         Up 30 seconds (healthy)
ErrorFix
pull access denied on gateway imagedocker logout and retry - a stale login to Docker Hub or ghcr.io can shadow the anonymous pull
could not select device driver "nvidia"NVIDIA Container Toolkit isn't installed, or Docker wasn't restarted after installing it
Gateway exits immediatelyParakeet container failed - check its logs

Step 3 - Test the Server

Generate a test audio file (macOS):

say -o test.aiff "Hello from my home server"

Or record a voice memo on your phone and AirDrop it over.

curl -X POST http://localhost:8080/v1/audio/transcriptions \
  -F "file=@test.aiff" \
  -F "model=parakeet-v3"
{"text":"Hello from my home server."}
# Check timing headers
curl -sS -D - -o /dev/null \
  -X POST http://localhost:8080/v1/audio/transcriptions \
  -F "file=@test.aiff" -F "model=parakeet-v3" | grep -i diction

That returns three headers, verified against a live gateway:

HeaderMeaning
X-Diction-Whisper-Msspeech model inference latency in milliseconds
X-Diction-Route-Modelwhich backend actually served the request
X-Diction-Route-Langdetected language, empty when detection didn't run

X-Diction-LLM-Ms is added when AI cleanup runs. X-Diction-Route-Model is the quickest way to confirm a model switch took effect, since an unrecognised model name silently falls back to DEFAULT_MODEL.

ResponseCause
Connection refusedGateway not running - docker compose ps
502 Bad GatewayParakeet unreachable, still loading, or DEFAULT_MODEL names a service that isn't running
400 Bad RequestUnsupported response_format (only json and text)
404 Not FoundURL typo - path must be exactly /v1/audio/transcriptions
OOM / container crashNot enough VRAM - Parakeet INT8 needs about 2 GB

Step 4 - Find Your Server's IP

macOS:

ipconfig getifaddr en0
# or
ifconfig | grep 'inet ' | grep -v 127.0.0.1

Linux:

hostname -I | awk '{print $1}'

Windows:

ipconfig | findstr IPv4

Pick the 192.168.x.x or 10.x.x.x address. Ignore anything starting with 100. - that's Tailscale.

Set a DHCP reservation in your router so the IP doesn't change on reboot. Or use Tailscale for a stable address that follows the machine anywhere.

Step 5 - Connect the App

Install Diction on your iPhone. On first launch:

  1. Settings → General → Keyboard → Keyboards → Add New Keyboard → Diction
  2. Tap Diction in the list → enable Allow Full Access
  3. Grant microphone access when prompted

Point it at your server:

  1. Open Diction → Preferences → Mode → Self-Hosted
  2. Enter your endpoint: http://192.168.1.42:8080 (your IP from Step 4)
  3. Tap Test connection - you should get a green check within a second

To dictate: open any app, tap a text field, long-press the globe icon (bottom-left of the iOS keyboard), pick Diction, tap the mic, speak, release.

Reach From Anywhere

Tailscale (recommended)

Tailscale creates a private WireGuard mesh between your devices. Install it on the server and iPhone, sign in to the same account, and use the 100.x.x.x Tailscale IP as your Diction endpoint. Works on cellular, café WiFi, anywhere. Free for personal use.

Cloudflare Tunnel (public URL, no port forwarding)

Add to your compose file:

  cloudflared:
    image: cloudflare/cloudflared:latest
    container_name: diction-cloudflared
    restart: unless-stopped
    command: tunnel --no-autoupdate run
    environment:
      TUNNEL_TOKEN: "${CLOUDFLARE_TUNNEL_TOKEN}"

Create a tunnel in the Cloudflare Zero Trust dashboard, grab the token, add it to .env, route the public hostname to http://gateway:8080. Free tier. Note: transcripts pass through Cloudflare's network (HTTPS-encrypted, but a third party is in the path).

ngrok (quick testing)

ngrok http 8080

Free tier URLs change on restart - good for a demo, not daily use.


No GPU? (CPU-only Whisper)

No NVIDIA GPU on the box you're using? Run Whisper on CPU instead. Slower than Parakeet, but it runs anywhere Docker does.

services:
  # The whisper server runs as uid 1000. A volume left behind by the older root-based
  # images is owned by root, so without this the server cannot write and every model
  # download fails with a permission error. No-op on a fresh install.
  whisper-models-init:
    image: dictionlabs/whisper-server:latest-cpu
    user: root
    volumes:
      - whisper-models:/cache
    entrypoint: ["sh", "-c", "chown -R 1000:1000 /cache"]
    restart: "no"

  whisper-small:
    image: dictionlabs/whisper-server:latest-cpu
    container_name: diction-whisper-small
    restart: unless-stopped
    depends_on:
      whisper-models-init:
        condition: service_completed_successfully
    volumes:
      - whisper-models:/home/ubuntu/.cache/huggingface/hub
    healthcheck:
      test: ["CMD", "curl", "-fsS", "-o", "/dev/null", "http://localhost:8000/health"]
      interval: 30s
      timeout: 5s
      retries: 3
      start_period: 120s

  # The whisper server never downloads a model on demand: it answers "not installed
  # locally" instead of fetching. Without this one-shot pull the stack starts cleanly
  # and then fails every transcription. Idempotent, exits once the model is on disk.
  whisper-small-pull:
    image: dictionlabs/whisper-server:latest-cpu
    depends_on:
      whisper-small:
        condition: service_healthy
    entrypoint: ["curl", "-fsS", "-X", "POST",
                 "http://whisper-small:8000/v1/models/DictionLabs/whisper-small-ct2"]
    restart: "no"

  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    depends_on:
      - whisper-small
    environment:
      DEFAULT_MODEL: small

volumes:
  whisper-models:

The whisper-models volume persists the model weights (~500 MB for small) so they survive container rebuilds. The first up takes a few minutes longer than later ones while whisper-small-pull downloads the model - it exits as soon as that finishes, and subsequent starts skip straight past it.

docker compose up -d
ErrorFix
exec format error on Apple SiliconEnable Rosetta in Docker Desktop → Settings → General
health: starting for > 3 minutesModel still downloading - docker compose logs -f whisper-small
Gateway exits immediatelyWhisper container failed - check its logs

Test it the same way as Step 3, with -F "model=small" instead of parakeet-v3. See Swap the Speech Model below for medium and large-v3-turbo.


Swap the Speech Model

Pick a compose profile and set DEFAULT_MODEL on the gateway to match:

DEFAULT_MODELCompose profileService nameWeights servedRAMNotes
smallsmallwhisper-smallDictionLabs/whisper-small-ct2~850 MBBest for CPU
mediummediumwhisper-mediumDictionLabs/whisper-medium-ct2~2.1 GBMore accurate, slower on CPU
large-v3-turbolargewhisper-large-turboDictionLabs/whisper-large-v3-turbo-ct2~2.3 GBHighest accuracy; slow on CPU, fast on GPU
parakeet-v3parakeetparakeetbaked into the image~2 GB25 European languages; fast on CPU, faster on GPU
docker compose --profile small up -d

DEFAULT_MODEL and the service name must both match the table: the gateway resolves backends by Docker hostname, so if the named service isn't running every request fails with 502.

An unrecognised model name is not an error. The gateway falls back to DEFAULT_MODEL, so a typo transcribes successfully with the wrong model rather than telling you. Check the X-Diction-Route-Model response header to see which backend actually served a request.

The weights column is informational. You do not set it anywhere: the gateway names the model when it forwards each request. The whisper server does not fetch models on demand, so the compose file ships a one-shot pull service per profile that installs the model before first use. Those are our own CTranslate2 builds of OpenAI's checkpoints, published at huggingface.co/DictionLabs, so the models your server pulls come from a namespace we control rather than a third party's conversion.

docker compose up -d   # recreates only the changed container

NVIDIA GPU: More Languages (large-v3-turbo)

Already set up Parakeet from Step 1? That covers 25 European languages. For the other 74, swap in Whisper large-v3-turbo running on GPU instead:

Parakeet TDT 0.6B v3Whisper Large-v3-turbo
WER (English)~6.3%7.4%
Latency (GPU)Sub-secondUnder 2s
VRAM~2 GB~2.3 GB
Languages25 European99
services:
  whisper-large-turbo:
    image: dictionlabs/whisper-server:latest-cuda
    container_name: diction-whisper-large-turbo
    restart: unless-stopped
    volumes:
      - whisper-models:/home/ubuntu/.cache/huggingface/hub
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    depends_on:
      - whisper-large-turbo
    environment:
      DEFAULT_MODEL: large-v3-turbo

volumes:
  whisper-models:

First boot downloads ~1.6 GB of model weights into the volume. Subsequent starts are instant.


Already Have a Voice Server?

Keep it. Use CUSTOM_BACKEND_URL to put the Diction Gateway in front of your existing server for WebSocket streaming and end-to-end encryption:

services:
  gateway:
    image: dictionlabs/gateway:latest
    platform: linux/amd64
    container_name: diction-gateway
    restart: unless-stopped
    ports:
      - "8080:8080"
    environment:
      CUSTOM_BACKEND_URL: http://your-existing-server:8000
      CUSTOM_BACKEND_MODEL: DictionLabs/whisper-small-ct2
VariableDescription
CUSTOM_BACKEND_AUTHAuthorization header forwarded to your backend, e.g. Bearer sk-xxx
CUSTOM_BACKEND_NEEDS_WAVSet to "true" if your backend only accepts WAV - the gateway converts with ffmpeg
CUSTOM_BACKEND_CANONICAL_IDHuggingFace-style ID advertised via /v1/models (default: CUSTOM_BACKEND_MODEL, or the literal custom if that is unset too)

AI Cleanup and Voice Editing (BYO LLM)

The gateway passes transcripts through any OpenAI-compatible LLM before returning them. You say "so um basically the meeting went well and uh they agreed to the timeline." The LLM returns "The meeting went well. They agreed to the timeline."

Enable the AI Companion toggle in the app. The gateway forwards the transcript to {LLM_BASE_URL}/chat/completions with your prompt, then returns the cleaned text. If the LLM fails, the raw transcript is returned -- dictation never breaks.

When LLM is configured, two additional text routes become available (requires TEXT_ROUTES_OPEN=true):

  • POST /v1/text/process?intent=edit -- apply a spoken voice instruction to text
  • POST /v1/text/process?intent=edit-selected -- apply an instruction to a selected text range
  • POST /v1/text/suggest -- return 2-3 alternative phrasings (always soft-fails)
  • POST /v1/text/summarize -- one-line summary of a saved note (always soft-fails)

Transcript cleanup also honours a "formatting" flag in the request context: when set (or omitted, which means on), LLM_PROMPT_FORMATTING is appended to the cleanup prompt so spoken lists and topic changes come back as line breaks and plain-text bullets.

These routes enable AI parity with Diction One on self-hosted gateways: voice-edit and suggestions work the same way, though results depend on the quality of your chosen model. Smaller models (under 7B) often do not follow edit instructions reliably -- 7B or larger is recommended for editing.

See AGENTS.md for the full wire format.

VariableRequiredDescription
LLM_BASE_URLYesOpenAI-compatible endpoint, e.g. https://api.openai.com/v1
LLM_MODELYesModel identifier, e.g. gpt-4o-mini
LLM_API_KEYNoBearer token. Not needed for local Ollama.
LLM_PROMPTNoSystem prompt for transcript cleanup. String or a file path starting with / (mount via volume). Defaults to the built-in cleanup prompt when empty.
LLM_PROMPT_EDITNoSystem prompt for voice-edit intent (?intent=edit). Defaults to the built-in edit prompt.
LLM_PROMPT_EDIT_SELECTEDNoSystem prompt for edit-selected intent. Defaults to the built-in edit-selected prompt.
LLM_PROMPT_SUGGESTNoSystem prompt for the suggest endpoint. Defaults to the built-in suggest prompt.
LLM_REASONING_EFFORTNoOpenAI-compatible reasoning effort such as none, low, medium, or high. Omitted by default.
LLM_PROMPT_FORMATTINGNoAppended to LLM_PROMPT when the client requests formatting. Defaults to the built-in formatting rules.
LLM_PROMPT_SUMMARYNoSystem prompt for /v1/text/summarize. Defaults to the built-in summary prompt.
TEXT_ROUTES_OPENNoSet to true to open /v1/text/process and /v1/text/suggest when AUTH_ENABLED=false. Default false (routes return 403 until explicitly opened).
DICTION_ENHANCE_TIMEOUT_MSNoHow long the cleanup pass may run while the user is waiting — /v1/audio/transcriptions, and /v1/audio/stream on voice-edit intents. Default 20000 (20s), sized for a local model on CPU. On timeout the raw transcript is returned, so nothing is ever lost. Set 0 for no limit.
DICTION_LIVE_ENHANCE_TIMEOUT_MSNoHow long the cleanup pass may run after the raw text has already been delivered, on /v1/audio/stream?split_enhance=true. Default 8000 (8s). Raising it much past 9s has no effect: the app stops waiting for the enhanced frame at 9s and keeps the raw text. Set 0 for no limit.

Both LLM_BASE_URL and LLM_MODEL must be set or the feature stays off.

Language hint: when the client's context carries a concrete language (not empty, not the auto-detect sentinel "auto"), the cleanup call gets a (Language: xx) hint appended to the user message — the built-in LLM_PROMPT tells the model to write in that language and correct wrong or missing accents/diacritics for it, never to translate. A custom LLM_PROMPT should say what to do with the hint if it matters to you; it is plain text appended after the transcript, not a template variable.

What the cleanup call sends. The transcript comes first, unlabelled, exactly as before. The descriptive context the user configured in the app then follows, one labelled line per item, each omitted entirely when unset — a user with none of them configured sends byte-for-byte the request earlier releases sent:

LineSourceCap
Custom words: a, b (also heard as: c)the user's My Words list50 entries
Tone: ...the user's Tone preset and About You description, mergednone
(Language: xx)the transcript's language, when concrete—

What it deliberately does NOT send, and why you should not add it. The client also supplies the text around the cursor, the earlier transcripts of the session, and the clipboard. The gateway decodes them and forwards none of them to the cleanup prompt. Measured against gpt-oss-20b with the built-in prompt: given a block of three earlier transcripts, the dictation "yeah that sounds good" came back as those three transcripts — the user's actual words gone from their text field. A long clipboard block was appended to the output verbatim. Custom words, tone and profile never did this.

The rule that survived: forward context that describes the speaker, never context that is itself prose the model could emit as the answer. The built-in prompt does say these lines are context and must not appear in the reply; a short prompt does not enforce it. If you write a custom LLM_PROMPT and want to feed it session or clipboard context, you own that trade — test it with a short dictation and a long context block before trusting it.

What the voice-edit call sends. For intent=edit (the user's cursor, nothing selected) the gateway sends Text: <before>‸<after> followed by Instruction: <what the user said>. The ‸ marks the cursor. The model is expected to return the full modified text with the marker removed; the gateway strips any ‸ the model leaves behind, because the app inserts the result verbatim. For intent=edit-selected the selection is the text and the surrounding context follows as Context before: / Context after: lines. An edit request with nothing to edit (no cursor context, or no selection) fails rather than guessing, so the app shows "Couldn't apply edit" instead of typing the user's spoken instruction into their document. A custom LLM_PROMPT_EDIT should account for the marker.

If cleanup keeps returning raw text, check the startup log line — it prints enhance_ms= and live_enhance_ms= — and time your model directly against LLM_BASE_URL. A local model slower than the budget is the usual cause; raise DICTION_ENHANCE_TIMEOUT_MS rather than switching cleanup off.

Behavior change from earlier releases: operators who set LLM_BASE_URL and LLM_MODEL without LLM_PROMPT now receive the built-in cleanup prompt automatically. Previously the gateway logged a warning and sent no system instructions. The default prompt is: "You are a transcript cleanup tool. Fix grammar, punctuation, and remove filler words. If a language is given, write in that language and correct wrong or missing accents or diacritics for it. Never translate. Lines labelled “Custom words” or “Tone” may follow the transcript: they describe the speaker, never part of what you return. Return only the corrected transcript, nothing else."

Option A - Cloud LLM (OpenAI, Groq, etc.)

echo "OPENAI_API_KEY=sk-your-key-here" > .env
  gateway:
    environment:
      DEFAULT_MODEL: small
      LLM_BASE_URL: "https://api.openai.com/v1"
      LLM_API_KEY: "${OPENAI_API_KEY}"
      LLM_MODEL: "gpt-4o-mini"
      LLM_PROMPT: "Clean up this voice transcription. Remove filler words (um, uh, like). Fix punctuation and capitalization. Return only the cleaned text, nothing else."

Docker Compose reads ${OPENAI_API_KEY} from .env automatically. Works with any OpenAI-compatible provider - Groq, Together, Fireworks, Mistral, OpenRouter - swap LLM_BASE_URL and LLM_MODEL.

Option B - Local Ollama (zero cost, fully private)

  ollama:
    image: ollama/ollama:latest
    container_name: diction-ollama
    restart: unless-stopped
    volumes:
      - ollama-models:/root/.ollama

  gateway:
    environment:
      DEFAULT_MODEL: small
      LLM_BASE_URL: "http://ollama:11434/v1"
      LLM_MODEL: "gemma2:9b"
      LLM_PROMPT: "Clean up this voice transcription. Remove filler words. Fix punctuation and capitalization. Return only the cleaned text, nothing else."

volumes:
  whisper-models:
  ollama-models:
docker compose up -d
docker exec diction-ollama ollama pull gemma2:9b
ModelMemoryNotes
gemma2:9b~6 GBBest cleanup quality at this size
qwen2.5:7b~5 GBStrong instruction following
llama3.1:8b~5 GBMost popular, well-tested
gemma3:4b~3 GBFor tighter machines

Models under 7B tend to answer questions about the transcript instead of cleaning it up. 7B or larger recommended.

Testing cleanup

curl -X POST "http://localhost:8080/v1/audio/transcriptions?enhance=true" \
  -F "file=@test.aiff" \
  -F "model=small"
# Confirm LLM fired - look for X-Diction-LLM-Ms in the output
curl -sS -D - -o /dev/null \
  -X POST "http://localhost:8080/v1/audio/transcriptions?enhance=true" \
  -F "file=@test.aiff" -F "model=small" | grep -i diction

Prompt file

Mount a file and point LLM_PROMPT at the path:

  gateway:
    volumes:
      - ./cleanup-prompt.txt:/config/prompt.txt:ro
    environment:
      LLM_PROMPT: "/config/prompt.txt"

If LLM_PROMPT starts with /, the gateway reads it as a file. Otherwise it uses the string directly.


NixOS

The repo ships a flake with a hardened systemd module - no Docker needed.

nix run github:DictionLabs/Diction#diction-gateway

Enable as a service:

{
  inputs.diction.url = "github:DictionLabs/Diction";

  outputs = { nixpkgs, diction, ... }: {
    nixosConfigurations.your-host = nixpkgs.lib.nixosSystem {
      modules = [
        diction.nixosModules.default
        {
          services.diction-gateway = {
            enable = true;
            openFirewall = true;
            # customBackend.url = "http://127.0.0.1:8000";
            # llm.baseUrl = "http://127.0.0.1:11434/v1";
            # llm.model = "gemma2:9b";
            # environmentFile = "/run/secrets/diction-gateway.env";
          };
        }
      ];
    };
  };
}

The unit runs under DynamicUser with ProtectSystem=strict, NoNewPrivileges, and a narrow syscall filter. Use environmentFile for secrets - they don't end up in the world-readable Nix store. Full option list: nix/module.nix.


OpenAI API Compatibility

The gateway implements the OpenAI audio transcription API - any client that works against api.openai.com/v1/audio/transcriptions works against a Diction gateway.

from openai import OpenAI

client = OpenAI(
    base_url="http://your-server:8080/v1",
    api_key="anything",  # not checked when AUTH_ENABLED=false
)

with open("audio.wav", "rb") as f:
    result = client.audio.transcriptions.create(
        file=f,
        model="small",            # or "DictionLabs/whisper-small-ct2"
        response_format="text",
    )
print(result)

Works with the Node SDK, LangChain, Flowise, n8n, or any tool that expects OpenAI's speech API.

Supported:

  • POST /v1/audio/transcriptions - file, model, language, prompt, response_format=json|text
  • GET /v1/models - returns an OpenAI-compatible data[] array plus a providers[] grouping consumed by the iOS app. HuggingFace IDs (DictionLabs/whisper-small-ct2, nvidia/parakeet-tdt-0.6b-v3) and short aliases (small, medium, large-v3-turbo, parakeet-v3) are both accepted. The older Systran/* and deepdml/* ids stay accepted as aliases, so existing scripts keep working.
  • WebSocket /v1/audio/stream - used by the Diction app for low-latency streaming

Not supported:

  • TTS (/v1/audio/speech)
  • response_format=verbose_json|srt|vtt (no word-level timestamps)
  • SSE streaming on REST (use WebSocket /v1/audio/stream instead)
  • Model download/delete (POST/DELETE /v1/models/{id})
  • OpenAI Realtime API (/v1/realtime)

Authentication is off by default (AUTH_ENABLED=false). Pass any non-empty string as the API key from the client - the gateway doesn't check it.

Do not set AUTH_ENABLED=true on a self-hosted deployment. It does not enable a shared secret. The middleware accepts only an Apple App Store JWS validated against the Apple root CA, or an HMAC trial token signed with TRIAL_SECRET. Both exist for Diction One. Turning it on locks you out of your own server. To expose a gateway publicly, put it behind a reverse proxy or tunnel that does the auth, or keep it on a VPN.

Error shape: errors return {"error":"<message>"}, except auth failures, which return {"error":"unauthorized","reason":"...","message":"..."}, not OpenAI's nested {"error":{"message":"...","type":"..."}}. Most SDKs surface these as HTTPError rather than APIError.


Privacy

  • On-device: Everything stays on your phone. No network connection is made.
  • Self-hosted: Audio goes to your server only. Neither the gateway nor the whisper server persists audio - it's transcribed and discarded.
  • AI cleanup enabled: The transcript (plain text, no audio) goes to your configured LLM. If you use Ollama locally, nothing leaves your machine.
  • Diction One (cloud): Audio is transcribed and immediately discarded. Not stored, not used for training.
  • Zero third-party SDKs in the app. No analytics, no tracking, no telemetry.
  • Full Access is required by iOS for any keyboard that makes network requests. What you type on the QWERTY keys never leaves your phone. The app sends your dictation audio, plus the text around your cursor when you edit by voice, and only to the endpoint you configured.

Read the full Privacy Policy.


Diction One

On-device and self-hosted are completely free with no word limits.

If you don't want to run a server, Diction One gives you a fine-tuned cloud model with advanced audio filtering - without the setup. Audio is sent to the Diction endpoint, transcribed, and immediately discarded. Pricing and trial details are in the app.


Contributing

Contributions are welcome. See CONTRIBUTING.md.

License

MIT. See LICENSE.

ios
keyboard
parakeet
speech-to-text
stt
voice
whisper
wispr
wispr-flow
wispr-flow-alternative
wisprflow-alternative

Significant stargazers

Igor Shubovych

101 followers · starred Jul 2026

Stéphane Busso

292 followers · starred Jul 2026