English | 繁體中文
A web app for listening to books on your phone with your eyes closed. The screen is one giant button: tap it and ask, and the AI answers out loud — about the plot, or a sentence you didn't understand — then carries on reading from exactly where it was interrupted.
.env; the thinking depth (effort) can be set too — see "Models and cost" belownpm install
cp .env.example .env # then put ANTHROPIC_API_KEY in .env (or use OpenAI: AI_PROVIDER=openai and OPENAI_API_KEY)
npm start # http://localhost:3000
No key yet and just want to see the interface: npm run start:mock (the AI answers are simulated; no API is called).
| System | How to start |
|---|---|
| Windows | Double-click start.bat; in a terminal type .\start.bat (PowerShell and recent cmd do not run files from the current folder by their bare name — typing just start.bat gives a "not recognized" error) |
| macOS / Linux | ./start.sh (the first time: chmod +x start.sh, or simply sh start.sh) |
Both scripts do the same thing: check that Node.js is installed and at least 20.12 (and tell you where to get it if not), run npm install the first time, then start the server; any extra arguments are passed straight to server.js (for example ./start.sh --mock).
"Port 3000 is already in use" at startup means the server is already running (in another window, or in the background) — just use it. To restart it, the message includes the commands to find and stop the old process on Windows and on macOS/Linux, or you can use another port with PORT=3100.
The project ships a .gitattributes that keeps .bat files as CRLF and .sh files as LF (a .sh converted to CRLF fails with "bad interpreter" on macOS/Linux); the sample books in fixtures/ are never touched byte for byte.
A phone's microphone only works over HTTPS, so http://192.168.x.x can't be used directly. There is a built-in LAN HTTPS mode:
HTTPS=1 to .env, then npm start (Windows: double-click start.bat; macOS/Linux: ./start.sh — see "Start scripts" above). The first start creates a small certificate authority (CA) that belongs only to you in .certs/, and issues a server certificate for this computer's current IP addresses.http://<computer-IP>:3001 (printed in the terminal) in the browser and follow the page to install the certificate once (iPhone: install the profile, then enable it under Settings → General → About → Certificate Trust Settings; Android: search Settings for "CA certificate" and install it).https://<computer-IP>:3000, and "Add to Home screen" to use it like an app. If the computer's IP changes, restarting the server issues a new certificate automatically — the phone does not need to reinstall anything.Security: the CA carries a name constraint, so it is only valid for private ranges (192.168.x.x, 10.x.x.x, 172.16–31.x.x, 100.64/10) and .local names — even if its private key leaked, it could not be used to forge an ordinary website. The private key stays in .certs/; never share that folder. If the computer is lost, remove the certificate in the phone's certificate settings. Anyone on the LAN can open the site and spend your API credit, so if guests use your Wi-Fi, set ACCESS_TOKEN in .env and enter the same password in the app's Settings.
Other options:
cloudflared tunnel --url http://localhost:3000 (without HTTPS=1) or Tailscale, which give you an https://… address.ACCESS_TOKEN.Android Chrome or iOS Safari is recommended (speech recognition and speech synthesis both depend on them).
| To do this | Do this |
|---|---|
| Ask a question | Press the giant button at the top (a rising chime ↑ and a vibration) and just speak; a short pause sends it automatically, or press again to send at once |
| Follow up during an answer | Press the giant button again while it is answering; the AI stops and listens for your next question |
| Cancel / skip the answer and go back to the book | Press the play button at the bottom middle (while answering it reads "Resume book") |
| Shake to ask (optional) | Turn on "Settings → Gestures → Shake the phone to start asking", then shake the phone hard back and forth — the same as pressing the giant button. Off by default |
| Pause / play | The big button at the bottom middle; the play/pause key on headphones works too |
| Back / forward | Left/right buttons = previous / next paragraph; long-press = previous / next chapter |
| Sleep timer | "Sleep" at the bottom: 15 / 30 / 60 minutes or "chapter end" |
| Control it by voice (next section) | When asking, just say "go back three sentences", "go to chapter five", "stop", "read faster"… the AI decides by itself whether it is a command or a question |
Notes on shake to ask:
Sounds: ↑ the microphone is open and you can speak (the screen said "Preparing the microphone…" before; the app opens the microphone first and only then chimes, so if you speak as soon as you hear it, no words are lost), ↓ question received, a soft tick = the AI is thinking, three rising notes = the answer starts, two low notes = cancelled or an error.
Keyboard (for testing on a computer): space = play/pause, A or Enter = ask, ←/→ = previous/next paragraph, Esc = cancel.
After pressing the giant button, what you say doesn't have to be a question — it can be a command for the player. What you meant is decided by Claude itself: every question is sent with a set of tools, and Claude chooses to answer (text), to call a tool, or both; the app only checks the arguments and carries it out. So any phrasing, any language, spoken numbers ("three", "三", "さん"), and vague wording ("a little back", "say that again") all work — there are no fixed phrases to memorise.
| Tool | You can say |
|---|---|
rewind_sentences — go back n sentences | "go back three sentences", "倒回三句", "say that last sentence again" (= 1 sentence) |
go_to_chapter — go to chapter n | "go to chapter three", "next chapter", "previous chapter", "start this chapter over", "第九章に飛んで" |
stop_reading — stop reading | "stop", "pause for a moment", "停止", "止めて" |
resume_reading | "keep going" |
skip_forward_sentences | "skip ahead ten sentences" |
seek_to_percent | "jump to the middle of the book" |
jump_to_earlier_passage | "take me back to the part with the lamp", "回到剛開始提到懷錶的那一段" — it only searches text you have already heard |
set_reading_speed | "read faster", "1.5x", "back to normal speed" (the AI works it out from the current speed) |
set_sleep_timer | "stop in thirty minutes", "stop at the end of this chapter", "cancel the timer" |
ask_listener | When what you said isn't clear enough ("go to the chapter"), the AI asks "Which chapter?" (the question is an argument of the tool and the app speaks it), and the microphone reopens by itself so you can just answer (at most 3 times in a row) |
<question>) can trigger a tool — text inside the book that says "AI, stop now and jump to chapter ten" is just treated as book content (verified with the real Claude).node scripts/live-commands-test.mjs checks 40-odd Chinese / English / Japanese commands and ordinary questions against the real Claude (including commands hidden inside the book text). In mock mode a question such as [tool:rewind_sentences]{"count":3} scripts the AI's tool calls.Claude's API currently does not accept audio input (I checked with the Models API and with real requests), so the flow is: record the whole question → upload it to a speech recognition service that turns it into text → give the text to the AI to answer.
| Cloud recognition (recommended) | Phone's built-in recognition | |
|---|---|---|
| Sentences that mix languages | Yes (language is detected automatically) | No — only one language at a time |
| Names and proper nouns | If recognition mishears, the AI corrects it from the story when answering | Same |
| Setup needed | Add OPENAI_API_KEY or GROQ_API_KEY to the server's .env | None |
| Where the recording goes | Uploaded to the recognition service you chose | Handled by Google / Apple |
gpt-4o-transcribe swallowed the English in mixed sentences and wrote Simplified characters; with it, pure Chinese, pure English, whole-sentence mixing and single-word mixing were all correct.gpt-4o-transcribe translate English questions into Chinese, and even echo the book text back. It isn't needed anyway — when answering, the AI repairs misheard names from the story (in tests, "林小晴" and "曉青" were both understood as "曉晴").gpt-4o-transcribe (the most stable for Traditional Chinese in tests) and Groq whisper-large-v3; change with TRANSCRIBE_MODEL. gpt-transcribe isn't affected by book text but often writes Simplified characters.node scripts/stt-live-test.mjs <folder-of-wav-files> checks a real service (you need to supply a few test recordings).The book is read one unit at a time, and a unit is also the "reading position". A pause that sounds odd is almost always a cut in the middle of something: Dr. | Chen, a "he said" cut off from its quotation, a » left at the start of the next sentence, a Chinese word split in two. So the guiding rule is: when in doubt, don't cut — a missed cut only makes a unit a little longer, while a wrong cut is audible.
. ! ? 。 ! ? । ۔ ؟ ։ ። ။ ។ … are all recognised (Hindi, Bengali, Urdu, Arabic, Armenian, Ethiopic, Burmese, Khmer…); the Greek question mark (which looks like a semicolon) is handled too.« French » (even the space before . »), „German“, »German«, 「Japanese」 and "English" always stays with its own sentence; a "he said", "dit-il", "—preguntó" or "と彼女は尋ねた" after the sentence end stays in the same sentence instead of being thrown to the start of the next one.public/js/abbreviations.js): multilingual abbreviations such as Dr. / Mrs. / z. B. / Sra. / ул. / TP. / Cad.; single letters (J. K., А. С.) and dotted forms (U.S., т.е., μ.μ.) by shape; No. 5 and Fig. 2 count as abbreviations only when a number follows (so No. He left. still cuts); etc. and the like, which often do end a sentence, are not treated as abbreviations; words that are also common ordinary words (art, sat, sun, ill, DNA…) are deliberately left out. Languages that write ordinals with a dot (German, Danish, Czech… — am 3. Mai) are handled according to the book's language."Really?" and she nodded, etc. and so on); dialogue dashes (— Как дела? — спросил Иван) likewise; things glued together such as 3.14, example.com/?x=1 and Wait...what are not cut either.Dr., quotation marks of many countries, and the full stops of Arabic, Hindi and others work in answers too.Older books are re-cut once, automatically: when the cutting rules change (SEG_VERSION), opening a book imported with the old rules makes the app put each paragraph back together, cut it again with the new rules, and carry your reading position and Q&A log over to the new units (converted by character proportion inside the paragraph, so you don't lose your place); the old AI memory is tied to the old units, so it is cleared and rebuilt as you listen. The result is identical to importing the book again (covered by tests).
Limitations: languages without full stops, such as Thai, can only be cut by spaces and length; a full stop after a single English letter (plan B., I.) is always treated as an abbreviation; the dot-ordinal rule depends on the book's language, so it doesn't apply if the language is detected wrongly; word boundaries for Chinese, Japanese, Thai and so on need browser support for Intl.Segmenter (all mainstream browsers have it now; without it the cut falls back to the length limit).
Principle: every single word the AI receives comes from before the "current position" (including the sentence being read). This is guaranteed by the structure of the code, not by asking the AI nicely.
book ──► speech units (one sentence per unit; the unit number = the reading position)
│
├─ summary tree (background) what you've already heard, summarized into a small note per ~2.8k tokens;
│ every 8 notes are merged into a coarser one (8→64→512…), so far-away parts
│ are coarse notes and nearby parts are fine notes
├─ full-text search (BM25) indexes only text *before* the position; the question + the last two
│ sentences find relevant older passages (no embeddings needed)
└─ verbatim window about 2.2k–6k tokens of exact text before the current position (when a
sentence is interrupted, the last one is here)
What one question sends to the AI = system rules + summaries (a few hundred to a few thousand tokens)
+ 3 relevant older passages + the verbatim window + the last few Q&As + the question
npm test checks the limit with a simulated novel of that length).public/js/context.js): after assembling, it checks that the largest unit number of everything is ≤ the current position, otherwise it throws an error and sends nothing. Summaries are only produced after that part has been heard; the search index only contains text before the position; and when you go back, old summaries and old Q&As from after the position are not used.The main model (answering questions, understanding voice commands, building the memory) is chosen in .env:
| Setting | Model | Needs |
|---|---|---|
(default) AI_PROVIDER=anthropic | Claude, default claude-sonnet-5-5 (for both answers and memory) | ANTHROPIC_API_KEY |
AI_PROVIDER=openai | OpenAI, default gpt-6.1-sol | OPENAI_API_KEY (the same key as cloud speech recognition) |
QA_PROVIDER for answers and SUMMARY_PROVIDER for the background memory (for example OpenAI for answers and Claude for the memory); whichever company is used needs its key, and the startup message and the Settings screen show which model is in use.QA_MODEL and SUMMARY_MODEL (for example gpt-5.5, gpt-5.4-mini, claude-opus-5-5…). The background summaries are "many small calls" — finishing a long novel can mean a few hundred; to save money set SUMMARY_MODEL to a cheaper model (claude-haiku-4-5 is roughly a quarter of the price, or gpt-5.4-mini); the summary quality is usually good enough.Effort (thinking depth): both providers share QA_EFFORT (answers) and SUMMARY_EFFORT (memory), low by default. Measured on gpt-6.1-sol (real question requests, time to the first streamed word):
| effort | First word | Notes |
|---|---|---|
low (default) | about 1.4 s | enough for spoken answers |
medium | about 1.4 s | hardly any extra time |
high | about 5 s | starts to feel like "thinking for a long time" |
xhigh | about 10 s | |
max | about 20 s | not recommended for spoken Q&A |
gpt-6.1-sol takes low–max (not none); gpt-5.5 also takes none and minimal; non-reasoning models such as gpt-4.1 take nothing (set off). When you don't set it, the server tries low first and, if refused, steps down automatically (medium → high → none) and remembers the result; a value you set is used as given, and if the model refuses it the API's error is shown directly (it lists the accepted values).AI_FALLBACKS=off turns them off.api.openai.com is used through the Responses API (gpt-6.1-sol can only use the "voice-control tools" together with effort there; Chat Completions simply refuses); any other OPENAI_BASE_URL (an OpenAI-compatible server) is used through Chat Completions; force one with OPENAI_API=responses|chat. Requests are always store: false, so OpenAI keeps no conversation.node scripts/live-test.mjs) and voice commands 42/42 (node scripts/live-commands-test.mjs; add AI_PROVIDER=openai to test OpenAI)..doc files are not supported.npm test
It covers sentence cutting (English, German, French, Spanish, Portuguese, Italian, Russian, Greek, Arabic, Hebrew, Hindi, Bengali, Urdu, Armenian, Ethiopic, Khmer, Burmese, Thai, Korean, Chinese and Japanese, plus the invariants "nothing lost, nothing too long, re-cutting never drifts" and the speed on a million characters), file structure, BM25, the summary tree, the no-spoiler invariant, the context-size limit, the exact requests the server sends to the Anthropic API (checked against a fake API: streaming, fallback retry, refusals and error mapping, access password, path traversal), and voice commands (argument checking, edge cases, the order of several commands, searching only text already heard, chapter titles not leaking, tool_use streaming).
node scripts/live-test.mjs runs an end-to-end check against the real Claude API using the key in .env (summarizes the whole sample book, asks questions at different positions, verifies there are no spoilers; costs a few cents).
python scripts/make-fixtures.py generates sample books in many formats in fixtures/ (EPUB/PDF/DOCX/ODT/FB2/RTF/MD/HTML/Big5 TXT…); with npm run start:mock they are available at /__fixtures/<filename>.
This project is released under CC0 1.0 (public domain dedication), see LICENSE: free to use, modify and distribute, including commercially, with no attribution needed.
CC0 covers only this project's own code and sample books. The packages installed by npm install (including pdf.js, JSZip and mammoth, which are copied into public/vendor/, as well as the Anthropic SDK and node-forge) have their own licenses, which CC0 does not affect.
server.js static files + /api/ask (SSE streaming) + /api/summarize; the API keys live only here
prompts.js system prompt and request assembly (no-spoiler rules, summary rules, the voice-command tool definitions)
openai.js OpenAI as the main model: turns requests into the Responses API or Chat Completions, streams replies, tool calls, effort negotiation
lan-cert.js LAN HTTPS: creates / renews the personal CA (with name constraints) and the server certificate
setup-page.js the page that explains installing the certificate on a phone (Chinese / English / Japanese, by Accept-Language)
start.bat start.sh Windows double-click start / macOS and Linux start script
public/js/
app.js UI and flow (the question flow: pause → listen → think → speak the answer → resume where it left off)
tts.js stt.js speech synthesis (sentence by sentence; "pause" = cancel and remember the position) / the phone's built-in speech recognition
recorder.js recording and "finished speaking" detection (for cloud recognition)
commands.js checks the tool calls the AI chose and turns them into player actions (pure functions, fully unit-tested)
i18n.js interface translations (Chinese / English / Japanese); test/i18n.test.js makes sure the three languages have exactly the same strings and no hard-coded Chinese
book.js segmenter.js abbreviations.js the book data structure, the sentence-cutting rules (multilingual), the abbreviation list
memory.js retrieval.js context.js the summary tree, BM25, and assembling the AI's context (the no-spoiler guarantee lives here)
parsers/ EPUB / PDF / DOCX / ODT / FB2 / RTF / HTML / MD / TXT (with encoding detection)
feedback.js session.js sounds and vibration / keeping the screen on and media keys
JavaScript
68.6%
Rich Text Format
18.2%
HTML
8.8%
CSS
2.1%
Python
1.9%
English | 繁體中文
A web app for listening to books on your phone with your eyes closed. The screen is one giant button: tap it and ask, and the AI answers out loud — about the plot, or a sentence you didn't understand — then carries on reading from exactly where it was interrupted.
.env; the thinking depth (effort) can be set too — see "Models and cost" belownpm install
cp .env.example .env # then put ANTHROPIC_API_KEY in .env (or use OpenAI: AI_PROVIDER=openai and OPENAI_API_KEY)
npm start # http://localhost:3000
No key yet and just want to see the interface: npm run start:mock (the AI answers are simulated; no API is called).
| System | How to start |
|---|---|
| Windows | Double-click start.bat; in a terminal type .\start.bat (PowerShell and recent cmd do not run files from the current folder by their bare name — typing just start.bat gives a "not recognized" error) |
| macOS / Linux | ./start.sh (the first time: chmod +x start.sh, or simply sh start.sh) |
Both scripts do the same thing: check that Node.js is installed and at least 20.12 (and tell you where to get it if not), run npm install the first time, then start the server; any extra arguments are passed straight to server.js (for example ./start.sh --mock).
"Port 3000 is already in use" at startup means the server is already running (in another window, or in the background) — just use it. To restart it, the message includes the commands to find and stop the old process on Windows and on macOS/Linux, or you can use another port with PORT=3100.
The project ships a .gitattributes that keeps .bat files as CRLF and .sh files as LF (a .sh converted to CRLF fails with "bad interpreter" on macOS/Linux); the sample books in fixtures/ are never touched byte for byte.
A phone's microphone only works over HTTPS, so http://192.168.x.x can't be used directly. There is a built-in LAN HTTPS mode:
HTTPS=1 to .env, then npm start (Windows: double-click start.bat; macOS/Linux: ./start.sh — see "Start scripts" above). The first start creates a small certificate authority (CA) that belongs only to you in .certs/, and issues a server certificate for this computer's current IP addresses.http://<computer-IP>:3001 (printed in the terminal) in the browser and follow the page to install the certificate once (iPhone: install the profile, then enable it under Settings → General → About → Certificate Trust Settings; Android: search Settings for "CA certificate" and install it).https://<computer-IP>:3000, and "Add to Home screen" to use it like an app. If the computer's IP changes, restarting the server issues a new certificate automatically — the phone does not need to reinstall anything.Security: the CA carries a name constraint, so it is only valid for private ranges (192.168.x.x, 10.x.x.x, 172.16–31.x.x, 100.64/10) and .local names — even if its private key leaked, it could not be used to forge an ordinary website. The private key stays in .certs/; never share that folder. If the computer is lost, remove the certificate in the phone's certificate settings. Anyone on the LAN can open the site and spend your API credit, so if guests use your Wi-Fi, set ACCESS_TOKEN in .env and enter the same password in the app's Settings.
Other options:
cloudflared tunnel --url http://localhost:3000 (without HTTPS=1) or Tailscale, which give you an https://… address.ACCESS_TOKEN.Android Chrome or iOS Safari is recommended (speech recognition and speech synthesis both depend on them).
| To do this | Do this |
|---|---|
| Ask a question | Press the giant button at the top (a rising chime ↑ and a vibration) and just speak; a short pause sends it automatically, or press again to send at once |
| Follow up during an answer | Press the giant button again while it is answering; the AI stops and listens for your next question |
| Cancel / skip the answer and go back to the book | Press the play button at the bottom middle (while answering it reads "Resume book") |
| Shake to ask (optional) | Turn on "Settings → Gestures → Shake the phone to start asking", then shake the phone hard back and forth — the same as pressing the giant button. Off by default |
| Pause / play | The big button at the bottom middle; the play/pause key on headphones works too |
| Back / forward | Left/right buttons = previous / next paragraph; long-press = previous / next chapter |
| Sleep timer | "Sleep" at the bottom: 15 / 30 / 60 minutes or "chapter end" |
| Control it by voice (next section) | When asking, just say "go back three sentences", "go to chapter five", "stop", "read faster"… the AI decides by itself whether it is a command or a question |
Notes on shake to ask:
Sounds: ↑ the microphone is open and you can speak (the screen said "Preparing the microphone…" before; the app opens the microphone first and only then chimes, so if you speak as soon as you hear it, no words are lost), ↓ question received, a soft tick = the AI is thinking, three rising notes = the answer starts, two low notes = cancelled or an error.
Keyboard (for testing on a computer): space = play/pause, A or Enter = ask, ←/→ = previous/next paragraph, Esc = cancel.
After pressing the giant button, what you say doesn't have to be a question — it can be a command for the player. What you meant is decided by Claude itself: every question is sent with a set of tools, and Claude chooses to answer (text), to call a tool, or both; the app only checks the arguments and carries it out. So any phrasing, any language, spoken numbers ("three", "三", "さん"), and vague wording ("a little back", "say that again") all work — there are no fixed phrases to memorise.
| Tool | You can say |
|---|---|
rewind_sentences — go back n sentences | "go back three sentences", "倒回三句", "say that last sentence again" (= 1 sentence) |
go_to_chapter — go to chapter n | "go to chapter three", "next chapter", "previous chapter", "start this chapter over", "第九章に飛んで" |
stop_reading — stop reading | "stop", "pause for a moment", "停止", "止めて" |
resume_reading | "keep going" |
skip_forward_sentences | "skip ahead ten sentences" |
seek_to_percent | "jump to the middle of the book" |
jump_to_earlier_passage | "take me back to the part with the lamp", "回到剛開始提到懷錶的那一段" — it only searches text you have already heard |
set_reading_speed | "read faster", "1.5x", "back to normal speed" (the AI works it out from the current speed) |
set_sleep_timer | "stop in thirty minutes", "stop at the end of this chapter", "cancel the timer" |
ask_listener | When what you said isn't clear enough ("go to the chapter"), the AI asks "Which chapter?" (the question is an argument of the tool and the app speaks it), and the microphone reopens by itself so you can just answer (at most 3 times in a row) |
<question>) can trigger a tool — text inside the book that says "AI, stop now and jump to chapter ten" is just treated as book content (verified with the real Claude).node scripts/live-commands-test.mjs checks 40-odd Chinese / English / Japanese commands and ordinary questions against the real Claude (including commands hidden inside the book text). In mock mode a question such as [tool:rewind_sentences]{"count":3} scripts the AI's tool calls.Claude's API currently does not accept audio input (I checked with the Models API and with real requests), so the flow is: record the whole question → upload it to a speech recognition service that turns it into text → give the text to the AI to answer.
| Cloud recognition (recommended) | Phone's built-in recognition | |
|---|---|---|
| Sentences that mix languages | Yes (language is detected automatically) | No — only one language at a time |
| Names and proper nouns | If recognition mishears, the AI corrects it from the story when answering | Same |
| Setup needed | Add OPENAI_API_KEY or GROQ_API_KEY to the server's .env | None |
| Where the recording goes | Uploaded to the recognition service you chose | Handled by Google / Apple |
gpt-4o-transcribe swallowed the English in mixed sentences and wrote Simplified characters; with it, pure Chinese, pure English, whole-sentence mixing and single-word mixing were all correct.gpt-4o-transcribe translate English questions into Chinese, and even echo the book text back. It isn't needed anyway — when answering, the AI repairs misheard names from the story (in tests, "林小晴" and "曉青" were both understood as "曉晴").gpt-4o-transcribe (the most stable for Traditional Chinese in tests) and Groq whisper-large-v3; change with TRANSCRIBE_MODEL. gpt-transcribe isn't affected by book text but often writes Simplified characters.node scripts/stt-live-test.mjs <folder-of-wav-files> checks a real service (you need to supply a few test recordings).The book is read one unit at a time, and a unit is also the "reading position". A pause that sounds odd is almost always a cut in the middle of something: Dr. | Chen, a "he said" cut off from its quotation, a » left at the start of the next sentence, a Chinese word split in two. So the guiding rule is: when in doubt, don't cut — a missed cut only makes a unit a little longer, while a wrong cut is audible.
. ! ? 。 ! ? । ۔ ؟ ։ ። ။ ។ … are all recognised (Hindi, Bengali, Urdu, Arabic, Armenian, Ethiopic, Burmese, Khmer…); the Greek question mark (which looks like a semicolon) is handled too.« French » (even the space before . »), „German“, »German«, 「Japanese」 and "English" always stays with its own sentence; a "he said", "dit-il", "—preguntó" or "と彼女は尋ねた" after the sentence end stays in the same sentence instead of being thrown to the start of the next one.public/js/abbreviations.js): multilingual abbreviations such as Dr. / Mrs. / z. B. / Sra. / ул. / TP. / Cad.; single letters (J. K., А. С.) and dotted forms (U.S., т.е., μ.μ.) by shape; No. 5 and Fig. 2 count as abbreviations only when a number follows (so No. He left. still cuts); etc. and the like, which often do end a sentence, are not treated as abbreviations; words that are also common ordinary words (art, sat, sun, ill, DNA…) are deliberately left out. Languages that write ordinals with a dot (German, Danish, Czech… — am 3. Mai) are handled according to the book's language."Really?" and she nodded, etc. and so on); dialogue dashes (— Как дела? — спросил Иван) likewise; things glued together such as 3.14, example.com/?x=1 and Wait...what are not cut either.Dr., quotation marks of many countries, and the full stops of Arabic, Hindi and others work in answers too.Older books are re-cut once, automatically: when the cutting rules change (SEG_VERSION), opening a book imported with the old rules makes the app put each paragraph back together, cut it again with the new rules, and carry your reading position and Q&A log over to the new units (converted by character proportion inside the paragraph, so you don't lose your place); the old AI memory is tied to the old units, so it is cleared and rebuilt as you listen. The result is identical to importing the book again (covered by tests).
Limitations: languages without full stops, such as Thai, can only be cut by spaces and length; a full stop after a single English letter (plan B., I.) is always treated as an abbreviation; the dot-ordinal rule depends on the book's language, so it doesn't apply if the language is detected wrongly; word boundaries for Chinese, Japanese, Thai and so on need browser support for Intl.Segmenter (all mainstream browsers have it now; without it the cut falls back to the length limit).
Principle: every single word the AI receives comes from before the "current position" (including the sentence being read). This is guaranteed by the structure of the code, not by asking the AI nicely.
book ──► speech units (one sentence per unit; the unit number = the reading position)
│
├─ summary tree (background) what you've already heard, summarized into a small note per ~2.8k tokens;
│ every 8 notes are merged into a coarser one (8→64→512…), so far-away parts
│ are coarse notes and nearby parts are fine notes
├─ full-text search (BM25) indexes only text *before* the position; the question + the last two
│ sentences find relevant older passages (no embeddings needed)
└─ verbatim window about 2.2k–6k tokens of exact text before the current position (when a
sentence is interrupted, the last one is here)
What one question sends to the AI = system rules + summaries (a few hundred to a few thousand tokens)
+ 3 relevant older passages + the verbatim window + the last few Q&As + the question
npm test checks the limit with a simulated novel of that length).public/js/context.js): after assembling, it checks that the largest unit number of everything is ≤ the current position, otherwise it throws an error and sends nothing. Summaries are only produced after that part has been heard; the search index only contains text before the position; and when you go back, old summaries and old Q&As from after the position are not used.The main model (answering questions, understanding voice commands, building the memory) is chosen in .env:
| Setting | Model | Needs |
|---|---|---|
(default) AI_PROVIDER=anthropic | Claude, default claude-sonnet-5-5 (for both answers and memory) | ANTHROPIC_API_KEY |
AI_PROVIDER=openai | OpenAI, default gpt-6.1-sol | OPENAI_API_KEY (the same key as cloud speech recognition) |
QA_PROVIDER for answers and SUMMARY_PROVIDER for the background memory (for example OpenAI for answers and Claude for the memory); whichever company is used needs its key, and the startup message and the Settings screen show which model is in use.QA_MODEL and SUMMARY_MODEL (for example gpt-5.5, gpt-5.4-mini, claude-opus-5-5…). The background summaries are "many small calls" — finishing a long novel can mean a few hundred; to save money set SUMMARY_MODEL to a cheaper model (claude-haiku-4-5 is roughly a quarter of the price, or gpt-5.4-mini); the summary quality is usually good enough.Effort (thinking depth): both providers share QA_EFFORT (answers) and SUMMARY_EFFORT (memory), low by default. Measured on gpt-6.1-sol (real question requests, time to the first streamed word):
| effort | First word | Notes |
|---|---|---|
low (default) | about 1.4 s | enough for spoken answers |
medium | about 1.4 s | hardly any extra time |
high | about 5 s | starts to feel like "thinking for a long time" |
xhigh | about 10 s | |
max | about 20 s | not recommended for spoken Q&A |
gpt-6.1-sol takes low–max (not none); gpt-5.5 also takes none and minimal; non-reasoning models such as gpt-4.1 take nothing (set off). When you don't set it, the server tries low first and, if refused, steps down automatically (medium → high → none) and remembers the result; a value you set is used as given, and if the model refuses it the API's error is shown directly (it lists the accepted values).AI_FALLBACKS=off turns them off.api.openai.com is used through the Responses API (gpt-6.1-sol can only use the "voice-control tools" together with effort there; Chat Completions simply refuses); any other OPENAI_BASE_URL (an OpenAI-compatible server) is used through Chat Completions; force one with OPENAI_API=responses|chat. Requests are always store: false, so OpenAI keeps no conversation.node scripts/live-test.mjs) and voice commands 42/42 (node scripts/live-commands-test.mjs; add AI_PROVIDER=openai to test OpenAI)..doc files are not supported.npm test
It covers sentence cutting (English, German, French, Spanish, Portuguese, Italian, Russian, Greek, Arabic, Hebrew, Hindi, Bengali, Urdu, Armenian, Ethiopic, Khmer, Burmese, Thai, Korean, Chinese and Japanese, plus the invariants "nothing lost, nothing too long, re-cutting never drifts" and the speed on a million characters), file structure, BM25, the summary tree, the no-spoiler invariant, the context-size limit, the exact requests the server sends to the Anthropic API (checked against a fake API: streaming, fallback retry, refusals and error mapping, access password, path traversal), and voice commands (argument checking, edge cases, the order of several commands, searching only text already heard, chapter titles not leaking, tool_use streaming).
node scripts/live-test.mjs runs an end-to-end check against the real Claude API using the key in .env (summarizes the whole sample book, asks questions at different positions, verifies there are no spoilers; costs a few cents).
python scripts/make-fixtures.py generates sample books in many formats in fixtures/ (EPUB/PDF/DOCX/ODT/FB2/RTF/MD/HTML/Big5 TXT…); with npm run start:mock they are available at /__fixtures/<filename>.
This project is released under CC0 1.0 (public domain dedication), see LICENSE: free to use, modify and distribute, including commercially, with no attribution needed.
CC0 covers only this project's own code and sample books. The packages installed by npm install (including pdf.js, JSZip and mammoth, which are copied into public/vendor/, as well as the Anthropic SDK and node-forge) have their own licenses, which CC0 does not affect.
server.js static files + /api/ask (SSE streaming) + /api/summarize; the API keys live only here
prompts.js system prompt and request assembly (no-spoiler rules, summary rules, the voice-command tool definitions)
openai.js OpenAI as the main model: turns requests into the Responses API or Chat Completions, streams replies, tool calls, effort negotiation
lan-cert.js LAN HTTPS: creates / renews the personal CA (with name constraints) and the server certificate
setup-page.js the page that explains installing the certificate on a phone (Chinese / English / Japanese, by Accept-Language)
start.bat start.sh Windows double-click start / macOS and Linux start script
public/js/
app.js UI and flow (the question flow: pause → listen → think → speak the answer → resume where it left off)
tts.js stt.js speech synthesis (sentence by sentence; "pause" = cancel and remember the position) / the phone's built-in speech recognition
recorder.js recording and "finished speaking" detection (for cloud recognition)
commands.js checks the tool calls the AI chose and turns them into player actions (pure functions, fully unit-tested)
i18n.js interface translations (Chinese / English / Japanese); test/i18n.test.js makes sure the three languages have exactly the same strings and no hard-coded Chinese
book.js segmenter.js abbreviations.js the book data structure, the sentence-cutting rules (multilingual), the abbreviation list
memory.js retrieval.js context.js the summary tree, BM25, and assembling the AI's context (the no-spoiler guarantee lives here)
parsers/ EPUB / PDF / DOCX / ODT / FB2 / RTF / HTML / MD / TXT (with encoding detection)
feedback.js session.js sounds and vibration / keeping the screen on and media keys
JavaScript
68.6%
Rich Text Format
18.2%
HTML
8.8%
CSS
2.1%
Python
1.9%