Tandem keeps your place in sync between an ebook and its audiobook. Self-hosted server, web reader/player, Android app, AI transcription for alignment.
Python
0
1,521 commits
updated Oct 4, 2026
Read a book, listen to the same book, and never lose your place. Tandem is a self-hosted server plus web and Android apps that keep one reading position across an ebook and its audiobook. It transcribes the audiobook with Whisper, aligns the transcript to the EPUB text sentence by sentence, and stores the position server-side — so you stop reading on the couch, carry on listening in the car, and open either format again exactly where you left off. It indexes books you already have on disk; nothing leaves your machine except the optional metadata lookups you switch on yourself.
Pending. No screenshots are committed yet, and this README will not fake them. The intended
set, to land under docs/images/:
| Planned file | Shows |
|---|---|
docs/images/library.png | Library grid with covers, pairing state and series |
docs/images/reader.png | EPUB reader with the switch-to-audio control |
docs/images/player.png | Audio player with the switch-to-ebook control |
docs/images/transcription-queue.png | Transcription queue mid-job |
Until then, run it and look — the quick start below brings the whole stack up in one command.
Pre-release: 0.1.0, no tagged releases yet. A single-maintainer project, built for and run on one self-hosted deployment. It is in daily use and carries a real test suite, so it is not a toy — but it has exactly one operator's worth of exposure, so expect rough edges on hardware and library layouts unlike theirs.
api_version handshake
(server/version.py) and the app warns on a mismatch, but nothing is
frozen before 1.0.| You need | Details |
|---|---|
| A server | x86-64 Linux host with Docker and Docker Compose. Postgres + FastAPI + nginx; no GPU needed for the server itself. Disk for your library, and a backup mount off the host disk. |
| Somewhere to transcribe | Either a CUDA GPU host running the Jetson worker (an Orin Nano 8 GB is the reference; medium is the default model), or a server image built with the local Whisper stack, which is multi-GB and slow on CPU. Transcription is optional — without it you get a library and two readers, but no cross-format sync. |
| A browser | Any current desktop or mobile browser. Installable as a PWA. |
| Android 8.0+ | API 26 or newer, for the Android app. Optional — the web app works on a phone. |
Read this before installing; most of it is by design and none of it is hidden.
.pdf files pair but never align; .mobi/.azw3 are indexed but must
be converted before they can be read or synced. See
docs/library-conventions.md.01.mp3 … 30.mp3 is deliberately skipped, not
imported — merge it to a single .m4b first.flowchart LR
W["Web app<br/>React + Vite"]
A["Android app<br/>Kotlin + Compose"]
subgraph host["Docker host"]
S["server<br/>FastAPI"]
D[("Postgres")]
end
EB[("ebooks<br/>volume")]
AB[("audiobooks<br/>volume")]
BK[("backups<br/>volume")]
J["Jetson worker<br/>faster-whisper"]
W -->|"REST + JWT"| S
A -->|"REST + JWT"| S
S <--> D
S -->|scan / stream| EB
S -->|scan / stream| AB
S -->|nightly dump| BK
S -.->|"transcription jobs, optional"| J
The server owns everything stateful: it scans the two library volumes, extracts metadata, pairs
ebooks with audiobooks, runs the transcription queue, generates sync maps, and is the only writer
of reading position. Clients are thin — they read and write position through the server
(PUT /api/sync/position/...) and hold only device-local caches.
| Path | Stack | What it is |
|---|---|---|
server/ | Python · FastAPI · SQLAlchemy (async) · Postgres | API, auth/RBAC, transcription queue, sync engine, import sources |
web/ | Vite · React 18 | Web app (responsive — same codebase serves desktop and mobile) |
android/ | Kotlin · Jetpack Compose | Android reader/listener client |
jetson/ | Python | Remote transcription worker (optional, separate host) |
Use Docker Compose for the full stack (server + Postgres + web). Copy the template once, then
customize your paths/secrets — docker-compose.yml is gitignored so your local copy never
conflicts with future pulls:
cp docker-compose.example.yml docker-compose.yml
cp .env.example .env
Then create the three secret files. They are not environment variables: anything in a container's
environment is printed by docker inspect and readable in /proc/1/environ, so the JWT signing
key, the Postgres password and the credential encryption keys live one-per-file under secrets/
and compose mounts them at /run/secrets/. secrets/ is gitignored; secrets/README.md explains
each file.
cd secrets
umask 077
python3 -c "import secrets; print(secrets.token_urlsafe(64))" > jwt_secret_key
python3 -c "import secrets; print(secrets.token_urlsafe(32))" > postgres_password
python3 -c 'from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())' > credential_enc_keys
cd ..
Edit docker-compose.yml: the ebook/audiobook/backup volume mounts, and the mandatory variables in
the table below. .env only carries PUID/PGID now — compose has to resolve those before any
container exists, so they cannot come from a secret. Both files are gitignored. Then:
docker compose up --build
The server applies its Alembic migrations on start (server/entrypoint.sh), so there's no separate
schema step. An existing database created before Alembic must be stamped once —
alembic stamp head — or upgrade head will try to create tables that already exist.
Run the
serverservice as a single process. Nouvicorn --workers, noWEB_CONCURRENCY, no second replica. The transcription queue claim, its cancellation state and the import/backup schedulers are all process-local, so a second worker transcribes the same audiobook twice and re-queues the other worker's running job. The server refuses to boot if the environment asks for more than one worker — see docs/operations.md, "Single process only".
Get the admin password. On a fresh database the server creates a single superadmin named
admin with a randomly generated password and writes it to the log — it is never admin:
docker compose logs server | grep -i superadmin
Log in and reset it. The account is flagged must_reset_password, so you are sent straight
to a forced password-reset screen. The flag is enforced by the server, not just the UI: until
the password is changed the API refuses every route except GET /auth/me,
POST /auth/change-password and POST /auth/logout with 403 password_reset_required. That
holds for the web app, the Android app and curl alike, so the temporary password cannot be used
for anything else.
If that password is ever lost, there is no self-service reset: an admin resets an ordinary
user from System → User Management, and a locked-out admin or sole superadmin is recovered with
docker compose exec server python -m scripts.reset_password <username> — see
docs/operations.md, "Account recovery".
Scan the library. System → Troubleshoot Library → Run Verification Scan walks the mounted ebook and audiobook directories, extracts metadata, and auto-pairs what it can match. See docs/library-conventions.md for the folder/filename patterns and the pairing rules.
Fix up the pairs. Anything ambiguous lands unpaired; pair it by hand under Pairs → Unpaired.
Queue transcription. Each pair needs a transcript before positions can be synced between the two formats. Queue it from the Transcription page, or turn on auto-transcribe to have new pairs queued automatically. See docs/transcription.md.
Set in the server service's environment: block. Everything the server reads lives in
server/config.py. PUID/PGID are the exception: they are compose
interpolation variables, read from a .env file next to docker-compose.yml (copy
.env.example) to fill the user: key, and the server process never sees them.
Secrets do not belong in environment:. Anything set there is printed by docker inspect and
sits in /proc/1/environ. Every secret below also has a <NAME>_FILE form that reads the value
from a file instead — that is what the template ships, backed by compose secrets:. See
docs/operations.md → Secrets.
Mandatory — with APP_ENV=prod (the default) the server refuses to start without these:
| Variable | Why |
|---|---|
JWT_SECRET_KEY | Signs access/refresh tokens. python -c "import secrets; print(secrets.token_urlsafe(64))" |
POSTGRES_PASSWORD | The booksync role's password. The server assembles DATABASE_URL from it (postgresql+asyncpg://booksync:<password>@db:5432/booksync) unless you set DATABASE_URL yourself; startup rejects the shipped booksync:booksync credentials either way. python -c "import secrets; print(secrets.token_urlsafe(32))" |
DATABASE_URL | Optional override of the assembled URL — set it only for an external database or non-default connection options |
CORS_ORIGINS | Comma-separated web origins. Startup rejects the wildcard *. |
CREDENTIAL_ENC_KEYS | Comma-separated Fernet keys encrypting import-source credentials (Audible auth blob, ABS token). First key encrypts, all are tried for decryption — rotation is "prepend a new key". python -c 'from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())' |
Set APP_ENV=dev for local development to allow the insecure zero-config defaults. Never in prod.
Secrets from files — the preferred form for all four of the above. <NAME>_FILE names a file
holding the value (one trailing newline is stripped); it wins over a plain <NAME> that is
also set, and a path that does not exist, cannot be read, or is empty is a startup error naming
the variable, never a silent fallback. docker-compose.example.yml wires these to compose
secrets: mounted at /run/secrets/; see docs/operations.md → Secrets
for how to create the files and how to migrate an existing install without changing any value.
| Variable | What it does |
|---|---|
JWT_SECRET_KEY_FILE | Path to a file holding JWT_SECRET_KEY |
CREDENTIAL_ENC_KEYS_FILE | Path to a file holding CREDENTIAL_ENC_KEYS |
POSTGRES_PASSWORD_FILE | Path to a file holding POSTGRES_PASSWORD. The same file is mounted into the db service, which reads this variable natively — one secret, both containers, guaranteed in sync |
DATABASE_URL_FILE | Path to a file holding a complete DATABASE_URL. Only needed for an external database whose URL differs from the assembled one |
GOOGLE_BOOKS_API_KEY_FILE | Path to a file holding GOOGLE_BOOKS_API_KEY |
ABS_API_TOKEN_FILE | Path to a file holding ABS_API_TOKEN |
Exposing the stack behind a reverse proxy? Set FORWARDED_ALLOW_IPS to your proxy's address,
prefer putting the proxy (e.g. Caddy) on the compose network pointed at server:8000, and bind
the published 8000/3000 ports to 127.0.0.1 or your LAN firewall so the proxy can't be
bypassed. Details: docs/operations.md → Reverse proxy.
Exposing it to the internet? Start from Caddyfile.example — it terminates
TLS and carries the security headers (HSTS, frame denial, a Content-Security-Policy) that nothing
inside the stack sets: docs/operations.md → Edge proxy.
Optional — all have working defaults:
| Variable | Default | What it does |
|---|---|---|
PUID / PGID | 1000 / 1000 | uid:gid the server and web containers run as (compose interpolation, from .env — not read by server/config.py). Set them to whatever owns your library on the host: stat -c '%u:%g' /path/to/your/ebooks. On an existing install, change the ownership of your app-data and backups directories before restarting — see docs/operations.md. db is unaffected; the postgres image manages its own user |
APP_ENV | prod | dev permits default secrets and wildcard CORS |
POSTGRES_USER / POSTGRES_DB | booksync / booksync | Role and database name used to assemble DATABASE_URL when it is not set explicitly. Must match the db service's own POSTGRES_USER/POSTGRES_DB |
POSTGRES_HOST / POSTGRES_PORT | db / 5432 | Where that assembled URL points — the compose service name and port |
EBOOK_DIR / AUDIOBOOK_DIR | /data/ebooks / /data/audiobooks | Library roots (mount your real folders here) |
APP_DATA_DIR / COVERS_DIR | /data/app / /data/app/covers | Extracted covers, working files |
IMPORTS_DIR | /data/imports | ACSM import staging (inbox / processed / failed) |
BACKUPS_DIR | /backups | Nightly + manual backups — mount this off the host disk |
TRANSCRIPTION_PROVIDER | remote_with_fallback | local, remote, or remote_with_fallback |
TRANSCRIPTION_REMOTE_URL | — | Jetson worker URL (port 9000) |
WHISPER_MODEL / WHISPER_DEVICE | medium / auto | Local Whisper model and device (auto/cpu/cuda) |
AUTO_TRANSCRIBE_ENABLED | false | Queue newly auto-matched pairs automatically |
ALLOW_PUBLIC_REGISTRATION | true | Kill switch only: false forces registration closed. Which of open / invite / closed a server is in otherwise is the registration_mode setting — see docs/operations.md, "Registration and invites" |
LOGIN_FAILURE_LIMIT / LOGIN_FAILURE_WINDOW_SECONDS | 10 / 900 | Failed logins per username before that username is 429'd for the rest of the window. Complements the per-IP limit; note that keying on the username lets anyone who knows one lock the real user out for the window, so don't lower it casually. See docs/operations.md |
FORWARDED_ALLOW_IPS | 127.0.0.1 (uvicorn default) | Reverse-proxy peer(s) whose X-Forwarded-For uvicorn trusts. Required behind any proxy — unset, all clients share the proxy's IP, so the login rate limit is one global bucket and audit logs record the proxy. Never *. See docs/operations.md |
GOOGLE_BOOKS_API_KEY | — | Raises the rate limit on the Google Books metadata search and the print page count lookup. Can also be set under System → Google Books, which wins over this |
ABS_URL / ABS_API_TOKEN / ABS_AUDIOBOOKS_PREFIX | — | Audiobookshelf metadata enrichment |
EXPENSIVE_READ_LIMIT | 30 | Per-user requests per window on the heavy read endpoints (troubleshoot issues, verify, disk usage, Calibre status); 429 with Retry-After past it (#208) |
EXPENSIVE_READ_WINDOW_SECONDS | 60 | Window for EXPENSIVE_READ_LIMIT |
SEARCH_READ_LIMIT | 60 | Per-user requests per window on library search and queue history |
SEARCH_READ_WINDOW_SECONDS | 60 | Window for SEARCH_READ_LIMIT |
EXTERNAL_METADATA_SEARCH_LIMIT | 20 | Per-user requests per window on external metadata search (spends the operator's provider quota) |
EXTERNAL_METADATA_SEARCH_WINDOW_SECONDS | 60 | Window for EXTERNAL_METADATA_SEARCH_LIMIT |
DISK_USAGE_CACHE_SECONDS | 300 | How long the disk-usage figure is cached |
CALIBRE_STATUS_CACHE_SECONDS | 300 | How long the Calibre availability check is cached |
CHAPTER_ENCODING_CACHE_SECONDS | 300 | How long a per-audiobook chapter-encoding check is cached (keyed on path, mtime, size) |
BACKUP_PROBE_CACHE_SECONDS | 60 | How long GET /api/health/backup caches its answer (#233) |
Tokens and auth throttles — sensible as they are; listed because they are settable, not because you should set them:
| Variable | Default | What it does |
|---|---|---|
JWT_ALGORITHM | HS256 | Signing algorithm for access/refresh tokens |
JWT_ACCESS_TOKEN_EXPIRE_MINUTES | 1440 | Access-token lifetime (24 h) |
JWT_REFRESH_TOKEN_EXPIRE_DAYS | 30 | Refresh-token lifetime. Refresh tokens are not rotated on use |
JWT_MEDIA_TOKEN_EXPIRE_MINUTES | 15 | Lifetime of the short-lived, resource-scoped token used for cover/audio URLs that cannot carry an Authorization header |
PASSWORD_CHANGE_FAILURE_LIMIT / PASSWORD_CHANGE_FAILURE_WINDOW_SECONDS | 5 / 900 | Failed current-password checks on /auth/change-password before that user is throttled. Keyed on the authenticated user id, so a low limit is safe here — unlike the login one |
REFRESH_FAILURE_LIMIT / REFRESH_FAILURE_WINDOW_SECONDS | 20 / 900 | Rejected /auth/refresh attempts per token subject before throttling. Only failures count |
Upload limits — layered with the proxy's own cap (see Caddyfile.example):
| Variable | Default | What it does |
|---|---|---|
UPLOAD_MAX_BYTES | 10737418240 (10 GiB) | Whole multipart body cap, refused with 413 before parsing |
MAX_UPLOAD_FILE_BYTES | 4294967296 (4 GiB) | Per-file cap enforced while the file streams to disk. Keep it ≤ UPLOAD_MAX_BYTES |
MAX_COVER_BYTES | 16777216 (16 MiB) | Per-file cap for cover images |
Sync tunables — contract-locked. These must stay equal to the clients' own constants; changing one here alone makes the ebook and the audiobook disagree about where you are. Change both clients too, and read docs/position-sync-contract.md first:
| Variable | Default | What it does |
|---|---|---|
DEFAULT_REWIND_SECONDS | 5 | How far back a text→audio handoff lands from the matched sentence (Android PlaybackOffsets.RESUME_REWIND_MS, web RESUME_REWIND_SECONDS) |
AUTO_COMPLETE_EPUB_PERCENT | 98.0 | Reading past this percentage marks the book finished — back matter means 100% is rarely reached |
AUTO_COMPLETE_AUDIO_TAIL_SECONDS | 120 | Listening to within this many seconds of the end marks the book finished |
The database wins over the environment for transcription settings. Provider, remote URL/key,
timeout, Whisper model, auto-transcribe, and the off-hours window are all editable from
System → Transcription Settings in the web UI, and the stored value is what the queue uses
(defaults in server/routers/settings.py). The env vars above are
the first-boot values.
The default server image is remote-transcription-only — it does not install the local
Whisper stack (torch + openai-whisper), whose CUDA wheels are multi-GB and slow to build.
Point transcription at the Jetson worker (TRANSCRIPTION_PROVIDER=remote; the default is
remote_with_fallback). To run Whisper on the server itself instead, build with the local stack:
docker compose build --build-arg INSTALL_LOCAL_WHISPER=1
(With a remote-only image, remote_with_fallback has no local fallback — it errors if the remote
is unreachable, so prefer TRANSCRIPTION_PROVIDER=remote unless you built with local Whisper.)
The Jetson Orin Nano worker deploys separately, on its own host — see jetson/README.md for the sparse-clone-and-deploy walkthrough.
With the default APP_ENV=prod the server serves no interactive schema: /docs, /redoc and
/openapi.json all return 404. That is deliberate — the published container port is normally
reachable on the LAN even when the reverse proxy forwards only /api/*, and the full route
inventory is reconnaissance rather than a feature. Set APP_ENV=dev (local development only, where
it also relaxes the secret/CORS startup checks) to get Swagger UI back at /docs.
The reference that is always available is docs/api.md, with the generated docs/openapi.json alongside it.
The server dumps Postgres and snapshots covers to /backups on a nightly schedule (point that
mount at a NAS path off the host disk). Create manual backups, restore, download, delete, and
configure the schedule/retention from System → Backups in the web UI.
Store CREDENTIAL_ENC_KEYS, JWT_SECRET_KEY, and POSTGRES_PASSWORD in your password
manager — without the Fernet keys, the encrypted import-source credentials in a restored dump
can't be decrypted. Full details, monitoring, and the restore/test-drill procedure:
docs/backup-restore.md.
Full index with one line per document: docs/README.md. Start there for the library conventions, transcription, Android and web-PWA guides, the operations and backup runbooks, the testing policy, and the position-sync contract. Deploying the remote transcription worker is jetson/README.md.
Privacy policy — what the apps send and to whom. Short version: Tandem is self-hosted, so the operator of the server you sign in to holds your data and we receive nothing; no ads, no analytics, no crash-reporting SDK. The one third party the Android app contacts is the dictionary service behind the reader's "Define" action, and only when you use it.
Tests run in CI on every PR and on every push to main. Write a failing test first, then make it pass.
Server: once per clone or worktree, cd server && ./setup-testenv.sh — it provisions the
exact interpreter CI uses (server/.python-version, currently 3.12) with uv. Then:
cd server && .venv/Scripts/python.exe -m pytest -q # .venv/bin/python on macOS/Linux
Runs on SQLite — no Docker needed. Use ptw for the auto-rerun TDD loop. Do not run the suite
with a global python: unpinned pytest and missing prod deps produce failures CI does not
have.
Web: cd web && npx vitest run (watch mode: npm test; the coverage gate is
npm run coverage).
Android: cd android && ./gradlew :app:testDebugUnitTest.
Jetson: cd jetson && python -m pytest test_server.py -v — runs in CI too, path-filtered on
jetson/**.
Coverage gates, which differ per platform: server and web have a global floor plus per-PR patch coverage (changed lines ≥80%); Android has only a Kover line-coverage floor — there is no patch gate; Jetson has none. Full policy, the fixtures/helpers available, and how to write a test: docs/testing.md.
Sync-matching logic exists on both server and Android and is pinned by shared golden vectors in
server/tests/fixtures/sync_parity/ — a change to one platform must update both.
Contributions should include tests for new/changed behavior — the PR template has the checklist. Read CONTRIBUTING.md before opening a PR (setup, the CI gates, and the changes that must touch more than one place); vulnerabilities go through SECURITY.md, never a public issue.
Notable changes land in CHANGELOG.md under Unreleased and are renamed to a
version heading when a release is cut. Cutting one — the version bump, the tag, the GitHub Release
and how to roll a deployment back — is docs/releasing.md. The Play Store route
for the Android app is a separate, longer process with its own document.
The entire repository is licensed under AGPL-3.0-only — see LICENSE. Under §13, anyone who runs a modified Tandem server for other users must offer those users the corresponding source code. The license choice follows the server's AGPL/GPL dependencies (ebooklib, mobi, audible, audible-cli).
The project used to be called BookSync and still answers to it in places that are expensive to
rename: the Android package com.booksync, the Postgres role/database booksync, the
booksync_db docker volume, booksync-db-*.dump backup files, and the repo slug Book-Sync.
Those are identifiers, not branding — leave them alone. Everything a user reads says Tandem. The
Android one is not merely expensive but permanent: Play binds an app's identity to the
applicationId of its first uploaded bundle, so once Tandem ships, com.booksync can never be
tidied to com.tandem — doing so would publish a second, unrelated app that no existing install
can update to.
Python
37.9%
Kotlin
37.3%
JavaScript
22.4%
CSS
2.1%
Tandem keeps your place in sync between an ebook and its audiobook. Self-hosted server, web reader/player, Android app, AI transcription for alignment.
Python
0
1,521 commits
updated Oct 4, 2026
Read a book, listen to the same book, and never lose your place. Tandem is a self-hosted server plus web and Android apps that keep one reading position across an ebook and its audiobook. It transcribes the audiobook with Whisper, aligns the transcript to the EPUB text sentence by sentence, and stores the position server-side — so you stop reading on the couch, carry on listening in the car, and open either format again exactly where you left off. It indexes books you already have on disk; nothing leaves your machine except the optional metadata lookups you switch on yourself.
Pending. No screenshots are committed yet, and this README will not fake them. The intended
set, to land under docs/images/:
| Planned file | Shows |
|---|---|
docs/images/library.png | Library grid with covers, pairing state and series |
docs/images/reader.png | EPUB reader with the switch-to-audio control |
docs/images/player.png | Audio player with the switch-to-ebook control |
docs/images/transcription-queue.png | Transcription queue mid-job |
Until then, run it and look — the quick start below brings the whole stack up in one command.
Pre-release: 0.1.0, no tagged releases yet. A single-maintainer project, built for and run on one self-hosted deployment. It is in daily use and carries a real test suite, so it is not a toy — but it has exactly one operator's worth of exposure, so expect rough edges on hardware and library layouts unlike theirs.
api_version handshake
(server/version.py) and the app warns on a mismatch, but nothing is
frozen before 1.0.| You need | Details |
|---|---|
| A server | x86-64 Linux host with Docker and Docker Compose. Postgres + FastAPI + nginx; no GPU needed for the server itself. Disk for your library, and a backup mount off the host disk. |
| Somewhere to transcribe | Either a CUDA GPU host running the Jetson worker (an Orin Nano 8 GB is the reference; medium is the default model), or a server image built with the local Whisper stack, which is multi-GB and slow on CPU. Transcription is optional — without it you get a library and two readers, but no cross-format sync. |
| A browser | Any current desktop or mobile browser. Installable as a PWA. |
| Android 8.0+ | API 26 or newer, for the Android app. Optional — the web app works on a phone. |
Read this before installing; most of it is by design and none of it is hidden.
.pdf files pair but never align; .mobi/.azw3 are indexed but must
be converted before they can be read or synced. See
docs/library-conventions.md.01.mp3 … 30.mp3 is deliberately skipped, not
imported — merge it to a single .m4b first.flowchart LR
W["Web app<br/>React + Vite"]
A["Android app<br/>Kotlin + Compose"]
subgraph host["Docker host"]
S["server<br/>FastAPI"]
D[("Postgres")]
end
EB[("ebooks<br/>volume")]
AB[("audiobooks<br/>volume")]
BK[("backups<br/>volume")]
J["Jetson worker<br/>faster-whisper"]
W -->|"REST + JWT"| S
A -->|"REST + JWT"| S
S <--> D
S -->|scan / stream| EB
S -->|scan / stream| AB
S -->|nightly dump| BK
S -.->|"transcription jobs, optional"| J
The server owns everything stateful: it scans the two library volumes, extracts metadata, pairs
ebooks with audiobooks, runs the transcription queue, generates sync maps, and is the only writer
of reading position. Clients are thin — they read and write position through the server
(PUT /api/sync/position/...) and hold only device-local caches.
| Path | Stack | What it is |
|---|---|---|
server/ | Python · FastAPI · SQLAlchemy (async) · Postgres | API, auth/RBAC, transcription queue, sync engine, import sources |
web/ | Vite · React 18 | Web app (responsive — same codebase serves desktop and mobile) |
android/ | Kotlin · Jetpack Compose | Android reader/listener client |
jetson/ | Python | Remote transcription worker (optional, separate host) |
Use Docker Compose for the full stack (server + Postgres + web). Copy the template once, then
customize your paths/secrets — docker-compose.yml is gitignored so your local copy never
conflicts with future pulls:
cp docker-compose.example.yml docker-compose.yml
cp .env.example .env
Then create the three secret files. They are not environment variables: anything in a container's
environment is printed by docker inspect and readable in /proc/1/environ, so the JWT signing
key, the Postgres password and the credential encryption keys live one-per-file under secrets/
and compose mounts them at /run/secrets/. secrets/ is gitignored; secrets/README.md explains
each file.
cd secrets
umask 077
python3 -c "import secrets; print(secrets.token_urlsafe(64))" > jwt_secret_key
python3 -c "import secrets; print(secrets.token_urlsafe(32))" > postgres_password
python3 -c 'from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())' > credential_enc_keys
cd ..
Edit docker-compose.yml: the ebook/audiobook/backup volume mounts, and the mandatory variables in
the table below. .env only carries PUID/PGID now — compose has to resolve those before any
container exists, so they cannot come from a secret. Both files are gitignored. Then:
docker compose up --build
The server applies its Alembic migrations on start (server/entrypoint.sh), so there's no separate
schema step. An existing database created before Alembic must be stamped once —
alembic stamp head — or upgrade head will try to create tables that already exist.
Run the
serverservice as a single process. Nouvicorn --workers, noWEB_CONCURRENCY, no second replica. The transcription queue claim, its cancellation state and the import/backup schedulers are all process-local, so a second worker transcribes the same audiobook twice and re-queues the other worker's running job. The server refuses to boot if the environment asks for more than one worker — see docs/operations.md, "Single process only".
Get the admin password. On a fresh database the server creates a single superadmin named
admin with a randomly generated password and writes it to the log — it is never admin:
docker compose logs server | grep -i superadmin
Log in and reset it. The account is flagged must_reset_password, so you are sent straight
to a forced password-reset screen. The flag is enforced by the server, not just the UI: until
the password is changed the API refuses every route except GET /auth/me,
POST /auth/change-password and POST /auth/logout with 403 password_reset_required. That
holds for the web app, the Android app and curl alike, so the temporary password cannot be used
for anything else.
If that password is ever lost, there is no self-service reset: an admin resets an ordinary
user from System → User Management, and a locked-out admin or sole superadmin is recovered with
docker compose exec server python -m scripts.reset_password <username> — see
docs/operations.md, "Account recovery".
Scan the library. System → Troubleshoot Library → Run Verification Scan walks the mounted ebook and audiobook directories, extracts metadata, and auto-pairs what it can match. See docs/library-conventions.md for the folder/filename patterns and the pairing rules.
Fix up the pairs. Anything ambiguous lands unpaired; pair it by hand under Pairs → Unpaired.
Queue transcription. Each pair needs a transcript before positions can be synced between the two formats. Queue it from the Transcription page, or turn on auto-transcribe to have new pairs queued automatically. See docs/transcription.md.
Set in the server service's environment: block. Everything the server reads lives in
server/config.py. PUID/PGID are the exception: they are compose
interpolation variables, read from a .env file next to docker-compose.yml (copy
.env.example) to fill the user: key, and the server process never sees them.
Secrets do not belong in environment:. Anything set there is printed by docker inspect and
sits in /proc/1/environ. Every secret below also has a <NAME>_FILE form that reads the value
from a file instead — that is what the template ships, backed by compose secrets:. See
docs/operations.md → Secrets.
Mandatory — with APP_ENV=prod (the default) the server refuses to start without these:
| Variable | Why |
|---|---|
JWT_SECRET_KEY | Signs access/refresh tokens. python -c "import secrets; print(secrets.token_urlsafe(64))" |
POSTGRES_PASSWORD | The booksync role's password. The server assembles DATABASE_URL from it (postgresql+asyncpg://booksync:<password>@db:5432/booksync) unless you set DATABASE_URL yourself; startup rejects the shipped booksync:booksync credentials either way. python -c "import secrets; print(secrets.token_urlsafe(32))" |
DATABASE_URL | Optional override of the assembled URL — set it only for an external database or non-default connection options |
CORS_ORIGINS | Comma-separated web origins. Startup rejects the wildcard *. |
CREDENTIAL_ENC_KEYS | Comma-separated Fernet keys encrypting import-source credentials (Audible auth blob, ABS token). First key encrypts, all are tried for decryption — rotation is "prepend a new key". python -c 'from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())' |
Set APP_ENV=dev for local development to allow the insecure zero-config defaults. Never in prod.
Secrets from files — the preferred form for all four of the above. <NAME>_FILE names a file
holding the value (one trailing newline is stripped); it wins over a plain <NAME> that is
also set, and a path that does not exist, cannot be read, or is empty is a startup error naming
the variable, never a silent fallback. docker-compose.example.yml wires these to compose
secrets: mounted at /run/secrets/; see docs/operations.md → Secrets
for how to create the files and how to migrate an existing install without changing any value.
| Variable | What it does |
|---|---|
JWT_SECRET_KEY_FILE | Path to a file holding JWT_SECRET_KEY |
CREDENTIAL_ENC_KEYS_FILE | Path to a file holding CREDENTIAL_ENC_KEYS |
POSTGRES_PASSWORD_FILE | Path to a file holding POSTGRES_PASSWORD. The same file is mounted into the db service, which reads this variable natively — one secret, both containers, guaranteed in sync |
DATABASE_URL_FILE | Path to a file holding a complete DATABASE_URL. Only needed for an external database whose URL differs from the assembled one |
GOOGLE_BOOKS_API_KEY_FILE | Path to a file holding GOOGLE_BOOKS_API_KEY |
ABS_API_TOKEN_FILE | Path to a file holding ABS_API_TOKEN |
Exposing the stack behind a reverse proxy? Set FORWARDED_ALLOW_IPS to your proxy's address,
prefer putting the proxy (e.g. Caddy) on the compose network pointed at server:8000, and bind
the published 8000/3000 ports to 127.0.0.1 or your LAN firewall so the proxy can't be
bypassed. Details: docs/operations.md → Reverse proxy.
Exposing it to the internet? Start from Caddyfile.example — it terminates
TLS and carries the security headers (HSTS, frame denial, a Content-Security-Policy) that nothing
inside the stack sets: docs/operations.md → Edge proxy.
Optional — all have working defaults:
| Variable | Default | What it does |
|---|---|---|
PUID / PGID | 1000 / 1000 | uid:gid the server and web containers run as (compose interpolation, from .env — not read by server/config.py). Set them to whatever owns your library on the host: stat -c '%u:%g' /path/to/your/ebooks. On an existing install, change the ownership of your app-data and backups directories before restarting — see docs/operations.md. db is unaffected; the postgres image manages its own user |
APP_ENV | prod | dev permits default secrets and wildcard CORS |
POSTGRES_USER / POSTGRES_DB | booksync / booksync | Role and database name used to assemble DATABASE_URL when it is not set explicitly. Must match the db service's own POSTGRES_USER/POSTGRES_DB |
POSTGRES_HOST / POSTGRES_PORT | db / 5432 | Where that assembled URL points — the compose service name and port |
EBOOK_DIR / AUDIOBOOK_DIR | /data/ebooks / /data/audiobooks | Library roots (mount your real folders here) |
APP_DATA_DIR / COVERS_DIR | /data/app / /data/app/covers | Extracted covers, working files |
IMPORTS_DIR | /data/imports | ACSM import staging (inbox / processed / failed) |
BACKUPS_DIR | /backups | Nightly + manual backups — mount this off the host disk |
TRANSCRIPTION_PROVIDER | remote_with_fallback | local, remote, or remote_with_fallback |
TRANSCRIPTION_REMOTE_URL | — | Jetson worker URL (port 9000) |
WHISPER_MODEL / WHISPER_DEVICE | medium / auto | Local Whisper model and device (auto/cpu/cuda) |
AUTO_TRANSCRIBE_ENABLED | false | Queue newly auto-matched pairs automatically |
ALLOW_PUBLIC_REGISTRATION | true | Kill switch only: false forces registration closed. Which of open / invite / closed a server is in otherwise is the registration_mode setting — see docs/operations.md, "Registration and invites" |
LOGIN_FAILURE_LIMIT / LOGIN_FAILURE_WINDOW_SECONDS | 10 / 900 | Failed logins per username before that username is 429'd for the rest of the window. Complements the per-IP limit; note that keying on the username lets anyone who knows one lock the real user out for the window, so don't lower it casually. See docs/operations.md |
FORWARDED_ALLOW_IPS | 127.0.0.1 (uvicorn default) | Reverse-proxy peer(s) whose X-Forwarded-For uvicorn trusts. Required behind any proxy — unset, all clients share the proxy's IP, so the login rate limit is one global bucket and audit logs record the proxy. Never *. See docs/operations.md |
GOOGLE_BOOKS_API_KEY | — | Raises the rate limit on the Google Books metadata search and the print page count lookup. Can also be set under System → Google Books, which wins over this |
ABS_URL / ABS_API_TOKEN / ABS_AUDIOBOOKS_PREFIX | — | Audiobookshelf metadata enrichment |
EXPENSIVE_READ_LIMIT | 30 | Per-user requests per window on the heavy read endpoints (troubleshoot issues, verify, disk usage, Calibre status); 429 with Retry-After past it (#208) |
EXPENSIVE_READ_WINDOW_SECONDS | 60 | Window for EXPENSIVE_READ_LIMIT |
SEARCH_READ_LIMIT | 60 | Per-user requests per window on library search and queue history |
SEARCH_READ_WINDOW_SECONDS | 60 | Window for SEARCH_READ_LIMIT |
EXTERNAL_METADATA_SEARCH_LIMIT | 20 | Per-user requests per window on external metadata search (spends the operator's provider quota) |
EXTERNAL_METADATA_SEARCH_WINDOW_SECONDS | 60 | Window for EXTERNAL_METADATA_SEARCH_LIMIT |
DISK_USAGE_CACHE_SECONDS | 300 | How long the disk-usage figure is cached |
CALIBRE_STATUS_CACHE_SECONDS | 300 | How long the Calibre availability check is cached |
CHAPTER_ENCODING_CACHE_SECONDS | 300 | How long a per-audiobook chapter-encoding check is cached (keyed on path, mtime, size) |
BACKUP_PROBE_CACHE_SECONDS | 60 | How long GET /api/health/backup caches its answer (#233) |
Tokens and auth throttles — sensible as they are; listed because they are settable, not because you should set them:
| Variable | Default | What it does |
|---|---|---|
JWT_ALGORITHM | HS256 | Signing algorithm for access/refresh tokens |
JWT_ACCESS_TOKEN_EXPIRE_MINUTES | 1440 | Access-token lifetime (24 h) |
JWT_REFRESH_TOKEN_EXPIRE_DAYS | 30 | Refresh-token lifetime. Refresh tokens are not rotated on use |
JWT_MEDIA_TOKEN_EXPIRE_MINUTES | 15 | Lifetime of the short-lived, resource-scoped token used for cover/audio URLs that cannot carry an Authorization header |
PASSWORD_CHANGE_FAILURE_LIMIT / PASSWORD_CHANGE_FAILURE_WINDOW_SECONDS | 5 / 900 | Failed current-password checks on /auth/change-password before that user is throttled. Keyed on the authenticated user id, so a low limit is safe here — unlike the login one |
REFRESH_FAILURE_LIMIT / REFRESH_FAILURE_WINDOW_SECONDS | 20 / 900 | Rejected /auth/refresh attempts per token subject before throttling. Only failures count |
Upload limits — layered with the proxy's own cap (see Caddyfile.example):
| Variable | Default | What it does |
|---|---|---|
UPLOAD_MAX_BYTES | 10737418240 (10 GiB) | Whole multipart body cap, refused with 413 before parsing |
MAX_UPLOAD_FILE_BYTES | 4294967296 (4 GiB) | Per-file cap enforced while the file streams to disk. Keep it ≤ UPLOAD_MAX_BYTES |
MAX_COVER_BYTES | 16777216 (16 MiB) | Per-file cap for cover images |
Sync tunables — contract-locked. These must stay equal to the clients' own constants; changing one here alone makes the ebook and the audiobook disagree about where you are. Change both clients too, and read docs/position-sync-contract.md first:
| Variable | Default | What it does |
|---|---|---|
DEFAULT_REWIND_SECONDS | 5 | How far back a text→audio handoff lands from the matched sentence (Android PlaybackOffsets.RESUME_REWIND_MS, web RESUME_REWIND_SECONDS) |
AUTO_COMPLETE_EPUB_PERCENT | 98.0 | Reading past this percentage marks the book finished — back matter means 100% is rarely reached |
AUTO_COMPLETE_AUDIO_TAIL_SECONDS | 120 | Listening to within this many seconds of the end marks the book finished |
The database wins over the environment for transcription settings. Provider, remote URL/key,
timeout, Whisper model, auto-transcribe, and the off-hours window are all editable from
System → Transcription Settings in the web UI, and the stored value is what the queue uses
(defaults in server/routers/settings.py). The env vars above are
the first-boot values.
The default server image is remote-transcription-only — it does not install the local
Whisper stack (torch + openai-whisper), whose CUDA wheels are multi-GB and slow to build.
Point transcription at the Jetson worker (TRANSCRIPTION_PROVIDER=remote; the default is
remote_with_fallback). To run Whisper on the server itself instead, build with the local stack:
docker compose build --build-arg INSTALL_LOCAL_WHISPER=1
(With a remote-only image, remote_with_fallback has no local fallback — it errors if the remote
is unreachable, so prefer TRANSCRIPTION_PROVIDER=remote unless you built with local Whisper.)
The Jetson Orin Nano worker deploys separately, on its own host — see jetson/README.md for the sparse-clone-and-deploy walkthrough.
With the default APP_ENV=prod the server serves no interactive schema: /docs, /redoc and
/openapi.json all return 404. That is deliberate — the published container port is normally
reachable on the LAN even when the reverse proxy forwards only /api/*, and the full route
inventory is reconnaissance rather than a feature. Set APP_ENV=dev (local development only, where
it also relaxes the secret/CORS startup checks) to get Swagger UI back at /docs.
The reference that is always available is docs/api.md, with the generated docs/openapi.json alongside it.
The server dumps Postgres and snapshots covers to /backups on a nightly schedule (point that
mount at a NAS path off the host disk). Create manual backups, restore, download, delete, and
configure the schedule/retention from System → Backups in the web UI.
Store CREDENTIAL_ENC_KEYS, JWT_SECRET_KEY, and POSTGRES_PASSWORD in your password
manager — without the Fernet keys, the encrypted import-source credentials in a restored dump
can't be decrypted. Full details, monitoring, and the restore/test-drill procedure:
docs/backup-restore.md.
Full index with one line per document: docs/README.md. Start there for the library conventions, transcription, Android and web-PWA guides, the operations and backup runbooks, the testing policy, and the position-sync contract. Deploying the remote transcription worker is jetson/README.md.
Privacy policy — what the apps send and to whom. Short version: Tandem is self-hosted, so the operator of the server you sign in to holds your data and we receive nothing; no ads, no analytics, no crash-reporting SDK. The one third party the Android app contacts is the dictionary service behind the reader's "Define" action, and only when you use it.
Tests run in CI on every PR and on every push to main. Write a failing test first, then make it pass.
Server: once per clone or worktree, cd server && ./setup-testenv.sh — it provisions the
exact interpreter CI uses (server/.python-version, currently 3.12) with uv. Then:
cd server && .venv/Scripts/python.exe -m pytest -q # .venv/bin/python on macOS/Linux
Runs on SQLite — no Docker needed. Use ptw for the auto-rerun TDD loop. Do not run the suite
with a global python: unpinned pytest and missing prod deps produce failures CI does not
have.
Web: cd web && npx vitest run (watch mode: npm test; the coverage gate is
npm run coverage).
Android: cd android && ./gradlew :app:testDebugUnitTest.
Jetson: cd jetson && python -m pytest test_server.py -v — runs in CI too, path-filtered on
jetson/**.
Coverage gates, which differ per platform: server and web have a global floor plus per-PR patch coverage (changed lines ≥80%); Android has only a Kover line-coverage floor — there is no patch gate; Jetson has none. Full policy, the fixtures/helpers available, and how to write a test: docs/testing.md.
Sync-matching logic exists on both server and Android and is pinned by shared golden vectors in
server/tests/fixtures/sync_parity/ — a change to one platform must update both.
Contributions should include tests for new/changed behavior — the PR template has the checklist. Read CONTRIBUTING.md before opening a PR (setup, the CI gates, and the changes that must touch more than one place); vulnerabilities go through SECURITY.md, never a public issue.
Notable changes land in CHANGELOG.md under Unreleased and are renamed to a
version heading when a release is cut. Cutting one — the version bump, the tag, the GitHub Release
and how to roll a deployment back — is docs/releasing.md. The Play Store route
for the Android app is a separate, longer process with its own document.
The entire repository is licensed under AGPL-3.0-only — see LICENSE. Under §13, anyone who runs a modified Tandem server for other users must offer those users the corresponding source code. The license choice follows the server's AGPL/GPL dependencies (ebooklib, mobi, audible, audible-cli).
The project used to be called BookSync and still answers to it in places that are expensive to
rename: the Android package com.booksync, the Postgres role/database booksync, the
booksync_db docker volume, booksync-db-*.dump backup files, and the repo slug Book-Sync.
Those are identifiers, not branding — leave them alone. Everything a user reads says Tandem. The
Android one is not merely expensive but permanent: Play binds an app's identity to the
applicationId of its first uploaded bundle, so once Tandem ships, com.booksync can never be
tidied to com.tandem — doing so would publish a second, unrelated app that no existing install
can update to.
Python
37.9%
Kotlin
37.3%
JavaScript
22.4%
CSS
2.1%