Super cool audio to text tool, one link to analyzed report application _ Developed By Kerwin
Python
22
6 commits
updated Sep 24, 2026
Free, local, unlimited-length transcription. Paste a YouTube/Bilibili link, get a transcript with timestamps. No account needed.
Built by Kapozux/Kerwin, student in Shanghai. In active development since March 2026.

Paste a link and Verbatim downloads the video, transcribes it, and writes a timestamped summary. Paste a whole channel or playlist instead, and it pulls every episode, transcribes each one, extracts evidence cards from every episode, and builds a portrait of the creator from those cards. If a cloud engine fails on a file, that file falls back to local Whisper automatically. It works with YouTube, Bilibili, and local audio/video files.
I can't sit through an hour-long video, but I still want to know what was said. Existing tools were either expensive or bad at Chinese, so I built my own. I've used it daily for six months and put 790+ hours of audio through it.
brew install ffmpeg yt-dlp
git clone https://github.com/Kapozux/_Verbatim_.git && cd _Verbatim_/getAudio
pip install -r requirements.txt
bash run.sh # open http://localhost:5001
Python 3.9+. Use Homebrew's yt-dlp, not the pip package.
| Engine | Good for |
|---|---|
| Whisper | Local and free (GPU-accelerated on Apple Silicon) |
| Gemini | Mixed Chinese/English |
| Gemini 3.5 | Tells speakers apart |
| QwenASR | Strong on Chinese |
| Precise | Hybrid: Gemini text + Alibaba speaker diarization |
Local Whisper transcription needs no API key. Summaries and creator analysis use Gemini; add keys in
Settings or in getAudio/.env:
GEMINI_API_KEY=... # Gemini engines, summaries, analysis
DASHSCOPE_API_KEY=... # QwenASR / Precise
OPENROUTER_API_KEY=... # optional: Claude as the analysis model
Every sentence in a creator analysis traces back to a specific timestamp in the original video. 33,832 evidence cards generated so far.
pipx install verbatim-transcribe-mcp
Claude Code and other agents can call Verbatim directly: transcribe links, search transcripts, run creator analysis.
| Lines of code | Commits | Transcripts | Hours transcribed | Evidence cards |
|---|---|---|---|---|
| ~16,000 | 100+ | 2,058 | 795 | 33,832 |
--cookies-from-browser requires the named browser to be installed locally.XHS_PROJECT env var) and uv. It's not a
pip dependency, since it drives a real logged-in browser session.Browser (SPA, SSE progress)
│ fetch / SSE
Flask (app.py)
├─ ThreadPoolExecutor + per-engine semaphores ── transcription concurrency
├─ global download / analysis semaphores ── throttling shared across all pipelines
├─ SQLite (taskdb.py, usage.db) ── task state, restart recovery, cost tracking
└─ results/<uuid>/…, results/_chains/<id>/… ── transcripts and creator analyses
harness.py gives two primitives, fanout
(concurrent, order-preserving map) and agent (one LLM call with retries and optional schema
validation). Every multi-step pipeline is a plain Python loop over them.sanitize.py collapses filler runs, dedups loops, repairs
timestamps, and removes silence-masked hallucinations, acting only when several signals agree.Key settings live in config.py (engine concurrency, Whisper model size, analysis model presets).
MIT
Transcription and analysis of third-party content is for private study. Respect the source platforms' terms and creators' rights.
Python
54.4%
JavaScript
28.7%
CSS
8.7%
HTML
8.0%
Super cool audio to text tool, one link to analyzed report application _ Developed By Kerwin
Python
22
6 commits
updated Sep 24, 2026
Free, local, unlimited-length transcription. Paste a YouTube/Bilibili link, get a transcript with timestamps. No account needed.
Built by Kapozux/Kerwin, student in Shanghai. In active development since March 2026.

Paste a link and Verbatim downloads the video, transcribes it, and writes a timestamped summary. Paste a whole channel or playlist instead, and it pulls every episode, transcribes each one, extracts evidence cards from every episode, and builds a portrait of the creator from those cards. If a cloud engine fails on a file, that file falls back to local Whisper automatically. It works with YouTube, Bilibili, and local audio/video files.
I can't sit through an hour-long video, but I still want to know what was said. Existing tools were either expensive or bad at Chinese, so I built my own. I've used it daily for six months and put 790+ hours of audio through it.
brew install ffmpeg yt-dlp
git clone https://github.com/Kapozux/_Verbatim_.git && cd _Verbatim_/getAudio
pip install -r requirements.txt
bash run.sh # open http://localhost:5001
Python 3.9+. Use Homebrew's yt-dlp, not the pip package.
| Engine | Good for |
|---|---|
| Whisper | Local and free (GPU-accelerated on Apple Silicon) |
| Gemini | Mixed Chinese/English |
| Gemini 3.5 | Tells speakers apart |
| QwenASR | Strong on Chinese |
| Precise | Hybrid: Gemini text + Alibaba speaker diarization |
Local Whisper transcription needs no API key. Summaries and creator analysis use Gemini; add keys in
Settings or in getAudio/.env:
GEMINI_API_KEY=... # Gemini engines, summaries, analysis
DASHSCOPE_API_KEY=... # QwenASR / Precise
OPENROUTER_API_KEY=... # optional: Claude as the analysis model
Every sentence in a creator analysis traces back to a specific timestamp in the original video. 33,832 evidence cards generated so far.
pipx install verbatim-transcribe-mcp
Claude Code and other agents can call Verbatim directly: transcribe links, search transcripts, run creator analysis.
| Lines of code | Commits | Transcripts | Hours transcribed | Evidence cards |
|---|---|---|---|---|
| ~16,000 | 100+ | 2,058 | 795 | 33,832 |
--cookies-from-browser requires the named browser to be installed locally.XHS_PROJECT env var) and uv. It's not a
pip dependency, since it drives a real logged-in browser session.Browser (SPA, SSE progress)
│ fetch / SSE
Flask (app.py)
├─ ThreadPoolExecutor + per-engine semaphores ── transcription concurrency
├─ global download / analysis semaphores ── throttling shared across all pipelines
├─ SQLite (taskdb.py, usage.db) ── task state, restart recovery, cost tracking
└─ results/<uuid>/…, results/_chains/<id>/… ── transcripts and creator analyses
harness.py gives two primitives, fanout
(concurrent, order-preserving map) and agent (one LLM call with retries and optional schema
validation). Every multi-step pipeline is a plain Python loop over them.sanitize.py collapses filler runs, dedups loops, repairs
timestamps, and removes silence-masked hallucinations, acting only when several signals agree.Key settings live in config.py (engine concurrency, Whisper model size, analysis model presets).
MIT
Transcription and analysis of third-party content is for private study. Respect the source platforms' terms and creators' rights.
Python
54.4%
JavaScript
28.7%
CSS
8.7%
HTML
8.0%