repo for speech to text using faster-whisper HF model
Create .venv and install dev/test toolchain (editable install, CPU torch wheels by default): ./run_uv.sh
GPU users: make use-gpu once per machine (writes .stt-variant.local), then ./run_uv.sh. make use-cpu to switch back; make show-variant to check. Mechanism: uv extras cpu / cu130 — see docs/Transcription_solution.md § "CPU / GPU install variants".
Run pre-commit via uv (uses the venv): uv run pre-commit run --all-files
Quick integration test run: make integration-local
Start Docker services (if needed): docker compose -f docker/docker-compose.yml up -d --build
Fallback to pip/venv:
Generate requirements.txt based on uv.lock: make export-reqs (writes requirements.txt for CPU and requirements-gpu.txt for the cu130 extra).
Batch audio transcription with Estonian (default) and English models, with speaker diarization on by default. Technical details → · Diarization setup (HF token) →
Windows (WSL): Copy a batch file from scripts/windows/ (e.g. transcribe_estonian_Desk.bat, transcribe_estonian_Teams.bat, transcribe_english_Desk.bat) into your audio folder and double-click. See scripts/windows/README.md for the full list and which variant each file targets.
Command line:
# Process folder (Estonian default)
.venv/bin/python scripts/transcribe_manager.py process /path/to/audio
# Show recent runs
.venv/bin/stt-faster db recent
# Use different model
.venv/bin/python scripts/transcribe_manager.py process /path/to/audio --preset large8gb
Model presets: et-large (Estonian, default), large8gb (English/multi, best accuracy), turbo (fast), distil (fastest)
Output: Files moved to processed/ subfolder with .txt transcripts (range timestamps + speaker labels by default). Use --output-format both for .json alongside, or --no-diarize to skip speaker labels. Failed files in failed/ subfolder.
Troubleshooting: For WSL paths use /mnt/c/Users/.... Delete ~/.local/share/stt-faster/runs.jsonl to reset run history.
Run stt-faster in Docker without installing Python or dependencies:
# Build production image (one time)
make docker-build-prod
# Process audio files using Docker wrapper
./scripts/transcribe-docker process /path/to/audio --preset turbo
# Or run directly
docker run --rm \
-v $(pwd):/workspace \
-v ~/.cache/hf:/home/appuser/.cache/hf \
stt-faster:latest process /workspace/audio --preset turbo
Features:
~/.cache/hf (persisted across runs)~/.local/share/stt-fasterSee: docker/README.md for full Docker documentation.
docs/AI_instructions.md for setup and usage.gemini.md quickstart; settings in .gemini/config.yaml and .gemini/settings.json for Gemini CLI and Gemini Code Assist..codex/ (use these configs in ~/.codex/) and AGENTS.md operational rules and guardrails for agents in this repo..cursor/ and .cursor/rules/*.mdc to guide Cursor behavior; .cursorignore for noise filtering.docs/ contains agent-focused references like AI_instructions.md, CODEX_RULES.md, and testing_approach.md.Pre-commit: Ruff (lint+format), Pyright (backend), Yamlfmt, Actionlint, Hadolint, Bandit, Detect-secrets. Run: uv run pre-commit run --all-files. Config: .pre-commit-config.yaml (+ .secrets.baseline).
Pre-push: Ruff format/check, Yamlfmt, Pyright, Unit tests; optional Semgrep + CodeQL via Act. Enable with make setup-hooks. Toggle via env: SKIP_LINT=1 SKIP_PYRIGHT=1 SKIP_TESTS=1 SKIP_LOCAL_SEC_SCANS=0.
GitHub CI (+local CI: Act cli):
python-lint-test.yml: Lint, Unit tests, Pyright. Integration/E2E run under Act (schedule/manual).meta-linters.yml: Actionlint, Yamlfmt, Hadolint on relevant changes.semgrep.yml, codeql.yml: Security scans on PR/schedule/manual.trivy_pip-audit.yml: pip-audit + Trivy on dep/Docker changes and schedule.Tests & coverage: tests/unit, tests/integration, tests/e2e. Fast path example: make unit. Coverage HTML: reports/coverage.
Logging: App logging in backend/config.py (level via LOG_LEVEL, Rich when TTY; optional file rotation via APP_LOG_DIR). Script logging helpers in scripts/common.sh. Unit tests cover both.
Use Makefile targets for common checks; see Makefile and docs/AI_instructions.md for details.
Run make pre-commit, which pins caches locally via UV_CACHE_DIR=./.uv-cache and PRE_COMMIT_HOME=./.pre-commit-cache.
If outbound network is unavailable and required wheels are not already cached, uv may fail (e.g., fetching filelock). Populate caches once in a networked environment or manually place the needed wheels under ./.uv-cache to reuse offline.
CI/Act Environment Alignment
.venv-ci) under act could drift from tools expecting .venv, leading to missing imports (e.g., dotenv, rich).pyrightconfig.json no longer sets venvPath/venv; Makefile passes --pythonpath so Pyright analyzes against the active interpreter.UV_PROJECT_ENVIRONMENT; uv creates/uses the in-project .venv by default.clean: false for actions/checkout so --bind doesn’t remove local files.Makefile targets listed above; interpreter selection and --pythonpath wiring..github/workflows/python-lint-test.yml environment no longer forces a venv name.MIT License - see LICENSE file.
273 commits
Python
87.4%
Batchfile
5.1%
Shell
3.7%
Makefile
2.4%
Dockerfile
1.3%
repo for speech to text using faster-whisper HF model
Create .venv and install dev/test toolchain (editable install, CPU torch wheels by default): ./run_uv.sh
GPU users: make use-gpu once per machine (writes .stt-variant.local), then ./run_uv.sh. make use-cpu to switch back; make show-variant to check. Mechanism: uv extras cpu / cu130 — see docs/Transcription_solution.md § "CPU / GPU install variants".
Run pre-commit via uv (uses the venv): uv run pre-commit run --all-files
Quick integration test run: make integration-local
Start Docker services (if needed): docker compose -f docker/docker-compose.yml up -d --build
Fallback to pip/venv:
Generate requirements.txt based on uv.lock: make export-reqs (writes requirements.txt for CPU and requirements-gpu.txt for the cu130 extra).
Batch audio transcription with Estonian (default) and English models, with speaker diarization on by default. Technical details → · Diarization setup (HF token) →
Windows (WSL): Copy a batch file from scripts/windows/ (e.g. transcribe_estonian_Desk.bat, transcribe_estonian_Teams.bat, transcribe_english_Desk.bat) into your audio folder and double-click. See scripts/windows/README.md for the full list and which variant each file targets.
Command line:
# Process folder (Estonian default)
.venv/bin/python scripts/transcribe_manager.py process /path/to/audio
# Show recent runs
.venv/bin/stt-faster db recent
# Use different model
.venv/bin/python scripts/transcribe_manager.py process /path/to/audio --preset large8gb
Model presets: et-large (Estonian, default), large8gb (English/multi, best accuracy), turbo (fast), distil (fastest)
Output: Files moved to processed/ subfolder with .txt transcripts (range timestamps + speaker labels by default). Use --output-format both for .json alongside, or --no-diarize to skip speaker labels. Failed files in failed/ subfolder.
Troubleshooting: For WSL paths use /mnt/c/Users/.... Delete ~/.local/share/stt-faster/runs.jsonl to reset run history.
Run stt-faster in Docker without installing Python or dependencies:
# Build production image (one time)
make docker-build-prod
# Process audio files using Docker wrapper
./scripts/transcribe-docker process /path/to/audio --preset turbo
# Or run directly
docker run --rm \
-v $(pwd):/workspace \
-v ~/.cache/hf:/home/appuser/.cache/hf \
stt-faster:latest process /workspace/audio --preset turbo
Features:
~/.cache/hf (persisted across runs)~/.local/share/stt-fasterSee: docker/README.md for full Docker documentation.
docs/AI_instructions.md for setup and usage.gemini.md quickstart; settings in .gemini/config.yaml and .gemini/settings.json for Gemini CLI and Gemini Code Assist..codex/ (use these configs in ~/.codex/) and AGENTS.md operational rules and guardrails for agents in this repo..cursor/ and .cursor/rules/*.mdc to guide Cursor behavior; .cursorignore for noise filtering.docs/ contains agent-focused references like AI_instructions.md, CODEX_RULES.md, and testing_approach.md.Pre-commit: Ruff (lint+format), Pyright (backend), Yamlfmt, Actionlint, Hadolint, Bandit, Detect-secrets. Run: uv run pre-commit run --all-files. Config: .pre-commit-config.yaml (+ .secrets.baseline).
Pre-push: Ruff format/check, Yamlfmt, Pyright, Unit tests; optional Semgrep + CodeQL via Act. Enable with make setup-hooks. Toggle via env: SKIP_LINT=1 SKIP_PYRIGHT=1 SKIP_TESTS=1 SKIP_LOCAL_SEC_SCANS=0.
GitHub CI (+local CI: Act cli):
python-lint-test.yml: Lint, Unit tests, Pyright. Integration/E2E run under Act (schedule/manual).meta-linters.yml: Actionlint, Yamlfmt, Hadolint on relevant changes.semgrep.yml, codeql.yml: Security scans on PR/schedule/manual.trivy_pip-audit.yml: pip-audit + Trivy on dep/Docker changes and schedule.Tests & coverage: tests/unit, tests/integration, tests/e2e. Fast path example: make unit. Coverage HTML: reports/coverage.
Logging: App logging in backend/config.py (level via LOG_LEVEL, Rich when TTY; optional file rotation via APP_LOG_DIR). Script logging helpers in scripts/common.sh. Unit tests cover both.
Use Makefile targets for common checks; see Makefile and docs/AI_instructions.md for details.
Run make pre-commit, which pins caches locally via UV_CACHE_DIR=./.uv-cache and PRE_COMMIT_HOME=./.pre-commit-cache.
If outbound network is unavailable and required wheels are not already cached, uv may fail (e.g., fetching filelock). Populate caches once in a networked environment or manually place the needed wheels under ./.uv-cache to reuse offline.
CI/Act Environment Alignment
.venv-ci) under act could drift from tools expecting .venv, leading to missing imports (e.g., dotenv, rich).pyrightconfig.json no longer sets venvPath/venv; Makefile passes --pythonpath so Pyright analyzes against the active interpreter.UV_PROJECT_ENVIRONMENT; uv creates/uses the in-project .venv by default.clean: false for actions/checkout so --bind doesn’t remove local files.Makefile targets listed above; interpreter selection and --pythonpath wiring..github/workflows/python-lint-test.yml environment no longer forces a venv name.MIT License - see LICENSE file.
273 commits
Python
87.4%
Batchfile
5.1%
Shell
3.7%
Makefile
2.4%
Dockerfile
1.3%