Yana/ft-llm-2026-ocr-dataset

Dataset

0

stars

3

commits

1

linked in READMEs

Apr 16, 2026

updated

document-understanding
finance
japanese
ocr
vlm

README

FT-LLM 2026 OCR Dataset

A Japanese document-OCR dataset used for Stage 1-1 caption + OCR pretraining of the COMPASS Vision-Language Model. Each sample pairs a rendered page image from a Japanese public-sector financial PDF (Cabinet Office, Financial Services Agency, Ministry of Finance) with its OCR-extracted markdown text. It is intended to teach the VLM's MLP projector to align vision tokens with Japanese text.

Part of the Compass collection.

License

Released under the Apache License 2.0.

Note on source materials and Japanese copyright law: Under Article 30-4 of the Japanese Copyright Act, the use of copyrighted works for the purpose of information analysis — including machine learning training — is a permitted use that does not require authorization from, or trigger license conditions of, the copyright holders. This dataset was produced in Japan on that basis, and the resulting artifacts are redistributed under Apache-2.0. Downstream users are responsible for complying with any applicable terms of the original source documents in their own jurisdiction.

Contributors

Yana

3 commits

Yana/ft-llm-2026-ocr-dataset

Dataset

0

stars

3

commits

1

linked in READMEs

Apr 16, 2026

updated

document-understanding
finance
japanese
ocr
vlm

README

FT-LLM 2026 OCR Dataset

A Japanese document-OCR dataset used for Stage 1-1 caption + OCR pretraining of the COMPASS Vision-Language Model. Each sample pairs a rendered page image from a Japanese public-sector financial PDF (Cabinet Office, Financial Services Agency, Ministry of Finance) with its OCR-extracted markdown text. It is intended to teach the VLM's MLP projector to align vision tokens with Japanese text.

Part of the Compass collection.

License

Released under the Apache License 2.0.

Note on source materials and Japanese copyright law: Under Article 30-4 of the Japanese Copyright Act, the use of copyrighted works for the purpose of information analysis — including machine learning training — is a permitted use that does not require authorization from, or trigger license conditions of, the copyright holders. This dataset was produced in Japan on that basis, and the resulting artifacts are redistributed under Apache-2.0. Downstream users are responsible for complying with any applicable terms of the original source documents in their own jurisdiction.

Contributors

Yana

3 commits