508-dev/journal-ocr

A journal OCR flow using local vision models.

Python

0

3 commits

updated Sep 14, 2026

See the code

README

journalocr

Handwritten journal photos → Markdown transcriptions, using local Ollama models.

images/<journal>/<page>.jpg ──► vision model A ──► work/<journal>/<page>/<A>.txt ──┐
                            └─► vision model B ──► work/<journal>/<page>/<B>.txt ──┤
                                                                                    ▼
                         out/<journal>/<page>.md ◄── reconciler LLM (+ glossary.md)

Each stage runs one model over every pending page before moving to the next model, so the ~20 GB models are loaded once per run rather than once per image.

Use

.venv/bin/python journalocr.py run                  # everything pending
.venv/bin/python journalocr.py run images/2026-03   # just one journal
.venv/bin/python journalocr.py run --stage reconcile --force   # after editing the reconcile prompt
.venv/bin/python journalocr.py status
.venv/bin/python journalocr.py score                # model accuracy on reviewed pages
.venv/bin/python journalocr.py export               # build dataset/

Models and options live in config.toml; prompts in prompts/. Rerunning is safe: existing outputs are skipped unless --force.

Reviewing

Every file in out/ starts with front matter:

image: images/2026-03/0042.jpg
image_sha256: 13fe2e…
status: draft

Correct the text, then set status: reviewed. The pipeline only ever overwrites pages whose status is draft, even with --force. Use any other value (in-review, …) to protect a page you're partway through.

When you fix a misread name, add it to glossary.md so the reconciler gets it right next time.

  • The path mirrors: images/X/Y.jpg ↔ out/X/Y.md ↔ work/X/Y/.
  • The front matter records the image path and its SHA-256, so export can find an image again by content if you rename or move it.
  • work/X/Y/draft.md is the pipeline's untouched output, kept so score can measure how far the pipeline was from your reviewed text.

export writes:

  • dataset/manifest.csv: every page with image, hash, transcription, draft, status
  • dataset/reviewed.jsonl: {"image", "image_sha256", "text"} for reviewed pages only, as training pairs

Don't edit or re-save the photos in images/. That changes the hash, and some editors apply the EXIF rotation tag (which is wrong on at least some of these photos).

508-dev/journal-ocr

A journal OCR flow using local vision models.

Python

0

3 commits

updated Sep 14, 2026

See the code

README

journalocr

Handwritten journal photos → Markdown transcriptions, using local Ollama models.

images/<journal>/<page>.jpg ──► vision model A ──► work/<journal>/<page>/<A>.txt ──┐
                            └─► vision model B ──► work/<journal>/<page>/<B>.txt ──┤
                                                                                    ▼
                         out/<journal>/<page>.md ◄── reconciler LLM (+ glossary.md)

Each stage runs one model over every pending page before moving to the next model, so the ~20 GB models are loaded once per run rather than once per image.

Use

.venv/bin/python journalocr.py run                  # everything pending
.venv/bin/python journalocr.py run images/2026-03   # just one journal
.venv/bin/python journalocr.py run --stage reconcile --force   # after editing the reconcile prompt
.venv/bin/python journalocr.py status
.venv/bin/python journalocr.py score                # model accuracy on reviewed pages
.venv/bin/python journalocr.py export               # build dataset/

Models and options live in config.toml; prompts in prompts/. Rerunning is safe: existing outputs are skipped unless --force.

Reviewing

Every file in out/ starts with front matter:

image: images/2026-03/0042.jpg
image_sha256: 13fe2e…
status: draft

Correct the text, then set status: reviewed. The pipeline only ever overwrites pages whose status is draft, even with --force. Use any other value (in-review, …) to protect a page you're partway through.

When you fix a misread name, add it to glossary.md so the reconciler gets it right next time.

  • The path mirrors: images/X/Y.jpg ↔ out/X/Y.md ↔ work/X/Y/.
  • The front matter records the image path and its SHA-256, so export can find an image again by content if you rename or move it.
  • work/X/Y/draft.md is the pipeline's untouched output, kept so score can measure how far the pipeline was from your reviewed text.

export writes:

  • dataset/manifest.csv: every page with image, hash, transcription, draft, status
  • dataset/reviewed.jsonl: {"image", "image_sha256", "text"} for reviewed pages only, as training pairs

Don't edit or re-save the photos in images/. That changes the hash, and some editors apply the EXIF rotation tag (which is wrong on at least some of these photos).

Languages

Python

100.0%