alex37529/voxspica

Free offline speech-to-text. Record or transcribe files with full privacy — no cloud, no tracking, works fully offline

4

29 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I made an offline speech-to-text app (VOSK, CPU-only, no cloud) — what should I build next? (r/SideProject)

I'm the developer of VoxSpica, an offline speech-to-text app for Windows. It runs on VOSK (Kaldi) on CPU — no cloud, no account, no subscription, no ads, and the audio never leaves the machine. What it does: * Live transcription from the microphone — text appears as you speak, with pause and stop *…

2

Oct 4, 2026

README

VoxSpica

VoxSpica

Free speech-to-text for Windows that never sends your voice anywhere.
Your voice stays on your computer. No account, no subscription, no cloud.

Download Guide Website

Windows License Interface languages Recognition languages Engine

English · Deutsch · Español · Français · Italiano · Русский · 简体中文


What it is

VoxSpica turns speech into text on Windows and nothing leaves the machine. The recognition engine is VOSK — a desktop build of Kaldi, open source — running on your CPU. No GPU, no network, no telemetry.

It works two ways:

  • Live dictation from a microphone, with the text appearing as you speak.
  • File transcription for anything the system can play — mp3, m4a, wav, and others if you have ffmpeg on the path.

Every recognition is stored in a local database you can search later.

Why offline

The dictation services you may have used send your audio to someone else's server. That is a reasonable design for them and a poor one for a tool whose job is to write down what you just said out loud — which is often a private document, a client's name, something you have not decided to publish yet.

Here the answer is structural rather than promised: the model is on your disk, the database is on your disk, and there is no code path that opens a socket.

Interface language selection
The interface speaks nine languages, each named in its own language.

Screens

Managing recognition models
Models — every language, its sizes, and what each one costs to download.
Choosing a microphone
Device — pick the microphone, with its real sample rate.
Recognition history
History — every recognition, searchable, deletable, copyable.
Interface language selection
Interface — switch the interface language at any time.

Get it

Per-language installerPortable archive

Seven installers, one per interface language. Each brings a ready recognition model for that language, so the first launch works with no internet at all. Pick your language and untick the box once you have the file — the links are straight below.

A single VoxSpica.exe in a zip. Nothing is installed, nothing is written to the system, and the folder can live on a USB stick. You pick the interface language and download a model on first run.

Portable builds are in Releases.

Scoop

If you already use Scoop:

scoop bucket add voxspica https://github.com/alex37529/voxspica-scoop
scoop install voxspica

That installs the portable archive and puts voxspica on your PATH. The manifest reads the checksum from the release's own SHA256SUMS.txt, so scoop update picks up a new version on its own.

The seven installers

Base URL: https://voxspica.4crytobot.xyz/downloads

The list names one version and is replaced when a new one is published. The portable archive is deliberately not in it — that one is in the release, and it is the one to take if you want an interface language that is not on this list. The flags mark the language, not the country: English is the British flag rather than the American one, because the interface is en and not en-US, and no installer ships a US-specific interface. They are served by flagcdn.com; if that host is unreachable the links below still work, only the pictures disappear.

Windows will warn you. The builds are not code-signed, so SmartScreen shows an unknown publisher. Click More info → Run anyway. See known limitations — signing is on the list.

First run

  1. Start the program. On the first launch it asks for the interface language.
  2. Choose the recognition language — the language you are going to speak. This is separate from the interface language and independent of it.
  3. Download a model for it. The small model is around 50 MB and is enough for drafts; the large one is more accurate, needs about 8 GB of RAM and takes 60–90 seconds to load the first time.
  4. Pick a microphone, press record, and talk.

The user guide walks through each screen, including what to do when a model is missing and how the models differ.

Languages

The interface is available in nine languages: Belarusian, German, English, Spanish, French, Italian, Russian, Ukrainian, Chinese.

Recognition works in 33 languages. They are listed with their model names in docs/LANGUAGES.md. Recognition and interface are independent — you can dictate Ukrainian into an English interface.

Command line

The program is not only a window. Everything it does is available from a console, which is what you want in a script or a hotkey launcher:

VoxSpica.exe mic --lang ru --size small     # dictate, print to stdout
VoxSpica.exe file meeting.mp3 --lang en-us  # transcribe a file
VoxSpica.exe devices                        # list microphones
VoxSpica.exe download --lang ru --size small
VoxSpica.exe list                           # languages and their models
VoxSpica.exe history --search "meeting"     # what was recognised before

What this does not do

Written plainly, because this is the part that usually gets discovered later.

  • Not code-signed. Windows shows a SmartScreen warning on every install.
  • No commas. VOSK outputs words without punctuation. VoxSpica restores capitalisation and puts a full stop at the end of a sentence, but it cannot place commas without parsing syntax — any simple rule produces "How, are you" instead of "How are you".
  • Windows x64 only. macOS and Linux are not built.
  • Models are separate. Nothing but the portable archive ships a model, and the models are 50 MB to 1.8 GB. The per-language installers are the exception: they carry one.

About this repository

This repository holds the releases only. The source code is not public and is not mirrored here — nothing in this repository contains it. The vX.Y.Z tags point at the small public commit that introduces this file, not at any application code. The files attached to each release are built binaries.

Website: https://voxspica.4crytobot.xyz

License

Proprietary, free of use, see LICENSE. Recognition is done by VOSK (Apache-2.0); recognition models are distributed by Alpha Cephei under the Apache-2.0 license as well.

speech-recognition
speech-recognition-api
speech-to-text
speech-to-text-app
voice-to-text

alex37529/voxspica

Free offline speech-to-text. Record or transcribe files with full privacy — no cloud, no tracking, works fully offline

4

29 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I made an offline speech-to-text app (VOSK, CPU-only, no cloud) — what should I build next? (r/SideProject)

I'm the developer of VoxSpica, an offline speech-to-text app for Windows. It runs on VOSK (Kaldi) on CPU — no cloud, no account, no subscription, no ads, and the audio never leaves the machine. What it does: * Live transcription from the microphone — text appears as you speak, with pause and stop *…

2

Oct 4, 2026

README

VoxSpica

VoxSpica

Free speech-to-text for Windows that never sends your voice anywhere.
Your voice stays on your computer. No account, no subscription, no cloud.

Download Guide Website

Windows License Interface languages Recognition languages Engine

English · Deutsch · Español · Français · Italiano · Русский · 简体中文


What it is

VoxSpica turns speech into text on Windows and nothing leaves the machine. The recognition engine is VOSK — a desktop build of Kaldi, open source — running on your CPU. No GPU, no network, no telemetry.

It works two ways:

  • Live dictation from a microphone, with the text appearing as you speak.
  • File transcription for anything the system can play — mp3, m4a, wav, and others if you have ffmpeg on the path.

Every recognition is stored in a local database you can search later.

Why offline

The dictation services you may have used send your audio to someone else's server. That is a reasonable design for them and a poor one for a tool whose job is to write down what you just said out loud — which is often a private document, a client's name, something you have not decided to publish yet.

Here the answer is structural rather than promised: the model is on your disk, the database is on your disk, and there is no code path that opens a socket.

Interface language selection
The interface speaks nine languages, each named in its own language.

Screens

Managing recognition models
Models — every language, its sizes, and what each one costs to download.
Choosing a microphone
Device — pick the microphone, with its real sample rate.
Recognition history
History — every recognition, searchable, deletable, copyable.
Interface language selection
Interface — switch the interface language at any time.

Get it

Per-language installerPortable archive

Seven installers, one per interface language. Each brings a ready recognition model for that language, so the first launch works with no internet at all. Pick your language and untick the box once you have the file — the links are straight below.

A single VoxSpica.exe in a zip. Nothing is installed, nothing is written to the system, and the folder can live on a USB stick. You pick the interface language and download a model on first run.

Portable builds are in Releases.

Scoop

If you already use Scoop:

scoop bucket add voxspica https://github.com/alex37529/voxspica-scoop
scoop install voxspica

That installs the portable archive and puts voxspica on your PATH. The manifest reads the checksum from the release's own SHA256SUMS.txt, so scoop update picks up a new version on its own.

The seven installers

Base URL: https://voxspica.4crytobot.xyz/downloads

The list names one version and is replaced when a new one is published. The portable archive is deliberately not in it — that one is in the release, and it is the one to take if you want an interface language that is not on this list. The flags mark the language, not the country: English is the British flag rather than the American one, because the interface is en and not en-US, and no installer ships a US-specific interface. They are served by flagcdn.com; if that host is unreachable the links below still work, only the pictures disappear.

Windows will warn you. The builds are not code-signed, so SmartScreen shows an unknown publisher. Click More info → Run anyway. See known limitations — signing is on the list.

First run

  1. Start the program. On the first launch it asks for the interface language.
  2. Choose the recognition language — the language you are going to speak. This is separate from the interface language and independent of it.
  3. Download a model for it. The small model is around 50 MB and is enough for drafts; the large one is more accurate, needs about 8 GB of RAM and takes 60–90 seconds to load the first time.
  4. Pick a microphone, press record, and talk.

The user guide walks through each screen, including what to do when a model is missing and how the models differ.

Languages

The interface is available in nine languages: Belarusian, German, English, Spanish, French, Italian, Russian, Ukrainian, Chinese.

Recognition works in 33 languages. They are listed with their model names in docs/LANGUAGES.md. Recognition and interface are independent — you can dictate Ukrainian into an English interface.

Command line

The program is not only a window. Everything it does is available from a console, which is what you want in a script or a hotkey launcher:

VoxSpica.exe mic --lang ru --size small     # dictate, print to stdout
VoxSpica.exe file meeting.mp3 --lang en-us  # transcribe a file
VoxSpica.exe devices                        # list microphones
VoxSpica.exe download --lang ru --size small
VoxSpica.exe list                           # languages and their models
VoxSpica.exe history --search "meeting"     # what was recognised before

What this does not do

Written plainly, because this is the part that usually gets discovered later.

  • Not code-signed. Windows shows a SmartScreen warning on every install.
  • No commas. VOSK outputs words without punctuation. VoxSpica restores capitalisation and puts a full stop at the end of a sentence, but it cannot place commas without parsing syntax — any simple rule produces "How, are you" instead of "How are you".
  • Windows x64 only. macOS and Linux are not built.
  • Models are separate. Nothing but the portable archive ships a model, and the models are 50 MB to 1.8 GB. The per-language installers are the exception: they carry one.

About this repository

This repository holds the releases only. The source code is not public and is not mirrored here — nothing in this repository contains it. The vX.Y.Z tags point at the small public commit that introduces this file, not at any application code. The files attached to each release are built binaries.

Website: https://voxspica.4crytobot.xyz

License

Proprietary, free of use, see LICENSE. Recognition is done by VOSK (Apache-2.0); recognition models are distributed by Alpha Cephei under the Apache-2.0 license as well.

speech-recognition
speech-recognition-api
speech-to-text
speech-to-text-app
voice-to-text