Yana/ft-llm-2026-qa-dataset

Dataset

1

stars

4

commits

1

linked in READMEs

Apr 16, 2026

updated

finance
instruction-tuning
japanese
vlm
vqa
Browse cluster: Japanese Vision-Language Models

README

FT-LLM 2026 QA Dataset

A Japanese visual-question-answering dataset used for Stage 1-2 visual instruction tuning of the COMPASS Vision-Language Model. Each sample contains a document or natural image together with one or more Japanese question–answer pairs, and is designed to give the VLM its instruction-following and VQA capabilities. Images are embedded in the dataset, so no external downloads are required.

Part of the Compass collection.

License

Released under the Apache License 2.0.

Note on source materials and Japanese copyright law: Under Article 30-4 of the Japanese Copyright Act, the use of copyrighted works for the purpose of information analysis — including machine learning training — is a permitted use that does not require authorization from, or trigger license conditions of, the copyright holders. This dataset was produced in Japan on that basis, and the resulting artifacts are redistributed under Apache-2.0. Downstream users are responsible for complying with any applicable terms of the original source documents in their own jurisdiction.

Contributors

Yana

4 commits

Yana/ft-llm-2026-qa-dataset

Dataset

1

stars

4

commits

1

linked in READMEs

Apr 16, 2026

updated

finance
instruction-tuning
japanese
vlm
vqa
Browse cluster: Japanese Vision-Language Models

README

FT-LLM 2026 QA Dataset

A Japanese visual-question-answering dataset used for Stage 1-2 visual instruction tuning of the COMPASS Vision-Language Model. Each sample contains a document or natural image together with one or more Japanese question–answer pairs, and is designed to give the VLM its instruction-following and VQA capabilities. Images are embedded in the dataset, so no external downloads are required.

Part of the Compass collection.

License

Released under the Apache License 2.0.

Note on source materials and Japanese copyright law: Under Article 30-4 of the Japanese Copyright Act, the use of copyrighted works for the purpose of information analysis — including machine learning training — is a permitted use that does not require authorization from, or trigger license conditions of, the copyright holders. This dataset was produced in Japan on that basis, and the resulting artifacts are redistributed under Apache-2.0. Downstream users are responsible for complying with any applicable terms of the original source documents in their own jurisdiction.

Contributors

Yana

4 commits