Free offline speech-to-text. Record or transcribe files with full privacy — no cloud, no tracking, works fully offline
See the code
Free speech-to-text for Windows that never sends your voice anywhere.
Your voice stays on your computer. No account, no subscription, no cloud.
English · Deutsch · Español · Français · Italiano · Русский · 简体中文
VoxSpica turns speech into text on Windows and nothing leaves the machine. The recognition engine is VOSK — a desktop build of Kaldi, open source — running on your CPU. No GPU, no network, no telemetry.
It works two ways:
Every recognition is stored in a local database you can search later.
The dictation services you may have used send your audio to someone else's server. That is a reasonable design for them and a poor one for a tool whose job is to write down what you just said out loud — which is often a private document, a client's name, something you have not decided to publish yet.
Here the answer is structural rather than promised: the model is on your disk, the database is on your disk, and there is no code path that opens a socket.
The interface speaks nine languages, each named in its own language.
![]() Models — every language, its sizes, and what each one costs to download. |
![]() Device — pick the microphone, with its real sample rate. |
![]() History — every recognition, searchable, deletable, copyable. |
![]() Interface — switch the interface language at any time. |
| Per-language installer | Portable archive |
|---|---|
|
Seven installers, one per interface language. Each brings a ready recognition model for that language, so the first launch works with no internet at all. Pick your language and untick the box once you have the file — the links are straight below. |
A single Portable builds are in Releases. |
If you already use Scoop:
scoop bucket add voxspica https://github.com/alex37529/voxspica-scoop
scoop install voxspica
That installs the portable archive and puts voxspica on your PATH. The
manifest reads the checksum from the release's own SHA256SUMS.txt, so
scoop update picks up a new version on its own.
Base URL: https://voxspica.4crytobot.xyz/downloads
VoxSpica-0.1.6-en-setup.exe — English, 119.3 MBVoxSpica-0.1.6-ru-setup.exe — Russian, 124.3 MBVoxSpica-0.1.6-de-setup.exe — German, 124.3 MBVoxSpica-0.1.6-fr-setup.exe — French, 120.6 MBVoxSpica-0.1.6-es-setup.exe — Spanish, 117.8 MBVoxSpica-0.1.6-it-setup.exe — Italian, 127.7 MBVoxSpica-0.1.6-zh-setup.exe — Chinese, 122.0 MBThe list names one version and is replaced when a new one is published. The
portable archive is deliberately not in it — that one is in the release, and it
is the one to take if you want an interface language that is not on this list.
The flags mark the language, not the country: English is the British flag rather
than the American one, because the interface is en and not en-US, and no
installer ships a US-specific interface. They are served by
flagcdn.com; if that host is unreachable the links below
still work, only the pictures disappear.
Windows will warn you. The builds are not code-signed, so SmartScreen shows an unknown publisher. Click More info → Run anyway. See known limitations — signing is on the list.
The user guide walks through each screen, including what to do when a model is missing and how the models differ.
The interface is available in nine languages: Belarusian, German, English, Spanish, French, Italian, Russian, Ukrainian, Chinese.
Recognition works in 33 languages. They are listed with their model names in docs/LANGUAGES.md. Recognition and interface are independent — you can dictate Ukrainian into an English interface.
The program is not only a window. Everything it does is available from a console, which is what you want in a script or a hotkey launcher:
VoxSpica.exe mic --lang ru --size small # dictate, print to stdout
VoxSpica.exe file meeting.mp3 --lang en-us # transcribe a file
VoxSpica.exe devices # list microphones
VoxSpica.exe download --lang ru --size small
VoxSpica.exe list # languages and their models
VoxSpica.exe history --search "meeting" # what was recognised before
Written plainly, because this is the part that usually gets discovered later.
This repository holds the releases only. The source code is not public and
is not mirrored here — nothing in this repository contains it. The vX.Y.Z tags
point at the small public commit that introduces this file, not at any
application code. The files attached to each release are built binaries.
Website: https://voxspica.4crytobot.xyz
Proprietary, free of use, see LICENSE. Recognition is done by VOSK (Apache-2.0); recognition models are distributed by Alpha Cephei under the Apache-2.0 license as well.
Free offline speech-to-text. Record or transcribe files with full privacy — no cloud, no tracking, works fully offline
See the code
Free speech-to-text for Windows that never sends your voice anywhere.
Your voice stays on your computer. No account, no subscription, no cloud.
English · Deutsch · Español · Français · Italiano · Русский · 简体中文
VoxSpica turns speech into text on Windows and nothing leaves the machine. The recognition engine is VOSK — a desktop build of Kaldi, open source — running on your CPU. No GPU, no network, no telemetry.
It works two ways:
Every recognition is stored in a local database you can search later.
The dictation services you may have used send your audio to someone else's server. That is a reasonable design for them and a poor one for a tool whose job is to write down what you just said out loud — which is often a private document, a client's name, something you have not decided to publish yet.
Here the answer is structural rather than promised: the model is on your disk, the database is on your disk, and there is no code path that opens a socket.
The interface speaks nine languages, each named in its own language.
![]() Models — every language, its sizes, and what each one costs to download. |
![]() Device — pick the microphone, with its real sample rate. |
![]() History — every recognition, searchable, deletable, copyable. |
![]() Interface — switch the interface language at any time. |
| Per-language installer | Portable archive |
|---|---|
|
Seven installers, one per interface language. Each brings a ready recognition model for that language, so the first launch works with no internet at all. Pick your language and untick the box once you have the file — the links are straight below. |
A single Portable builds are in Releases. |
If you already use Scoop:
scoop bucket add voxspica https://github.com/alex37529/voxspica-scoop
scoop install voxspica
That installs the portable archive and puts voxspica on your PATH. The
manifest reads the checksum from the release's own SHA256SUMS.txt, so
scoop update picks up a new version on its own.
Base URL: https://voxspica.4crytobot.xyz/downloads
VoxSpica-0.1.6-en-setup.exe — English, 119.3 MBVoxSpica-0.1.6-ru-setup.exe — Russian, 124.3 MBVoxSpica-0.1.6-de-setup.exe — German, 124.3 MBVoxSpica-0.1.6-fr-setup.exe — French, 120.6 MBVoxSpica-0.1.6-es-setup.exe — Spanish, 117.8 MBVoxSpica-0.1.6-it-setup.exe — Italian, 127.7 MBVoxSpica-0.1.6-zh-setup.exe — Chinese, 122.0 MBThe list names one version and is replaced when a new one is published. The
portable archive is deliberately not in it — that one is in the release, and it
is the one to take if you want an interface language that is not on this list.
The flags mark the language, not the country: English is the British flag rather
than the American one, because the interface is en and not en-US, and no
installer ships a US-specific interface. They are served by
flagcdn.com; if that host is unreachable the links below
still work, only the pictures disappear.
Windows will warn you. The builds are not code-signed, so SmartScreen shows an unknown publisher. Click More info → Run anyway. See known limitations — signing is on the list.
The user guide walks through each screen, including what to do when a model is missing and how the models differ.
The interface is available in nine languages: Belarusian, German, English, Spanish, French, Italian, Russian, Ukrainian, Chinese.
Recognition works in 33 languages. They are listed with their model names in docs/LANGUAGES.md. Recognition and interface are independent — you can dictate Ukrainian into an English interface.
The program is not only a window. Everything it does is available from a console, which is what you want in a script or a hotkey launcher:
VoxSpica.exe mic --lang ru --size small # dictate, print to stdout
VoxSpica.exe file meeting.mp3 --lang en-us # transcribe a file
VoxSpica.exe devices # list microphones
VoxSpica.exe download --lang ru --size small
VoxSpica.exe list # languages and their models
VoxSpica.exe history --search "meeting" # what was recognised before
Written plainly, because this is the part that usually gets discovered later.
This repository holds the releases only. The source code is not public and
is not mirrored here — nothing in this repository contains it. The vX.Y.Z tags
point at the small public commit that introduces this file, not at any
application code. The files attached to each release are built binaries.
Website: https://voxspica.4crytobot.xyz
Proprietary, free of use, see LICENSE. Recognition is done by VOSK (Apache-2.0); recognition models are distributed by Alpha Cephei under the Apache-2.0 license as well.