tokyotech-llm/lmsys-chat-1m-synth

Dataset

23

stars

76

commits

Feb 20, 2026

updated

README

LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M

This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M.

Additional Materials

We distribute Python scripts used to develop the dataset under the ./materials/ directory. The directory includes scripts for generating assistant responses and scoring preferences. These scripts are provided as-is solely for reproducibility of research purpose. We do not support or take responsibility in using these scripts.

License Information - Dataset

We publish the synthesized portion of the dataset under mixed licenses for each subset as follows:

User Instructions Translated into Japanese

The subsets lmsys-chat-1m-first-turn-user-instructions-ja.jsonl.gz.gpg and lmsys-chat-1m-first-turn-user-instructions-ja+unsafe.jsonl.gz.gpg, termed "Japanese Instructions," are distributed under the LMSYS-Chat-1M Dataset License Agreement. To access the original dataset and obtain the decryption key for the Japanese Instructions, you must agree to the license and provide your contact information. Please note that the "Right to Request Deletion" clause from the LMSYS-Chat-1M Dataset License also applies to Japanese Instructions: The original dataset authors retain the right to request you to delete all copies of the Japanese Instructions (in whole or in part) in your possession and control. You are required to comply with any and all such requests.

Assistant Responses and Preference Scores

The subset llama3.1-lmsys-chat-1m-synth-ja+en.jsonl.gz is distributed under the LLAMA 3.1 COMMUNITY LICENSE AGREEMENT.

The subsets gemma2-lmsys-chat-1m-synth-ja+en.jsonl.gz and gemma3-lmsys-chat-1m-synth-ja+en.jsonl.gz are distributed under the GEMMA TERMS OF USE.

The subset gpt-oss-lmsys-chat-1m-synth-ja+en.jsonl.gz is distributed under the Apache License, Version 2.0, in accordance with the gpt-oss license.

License Information - Scripts

We distribute the Python and Shell scripts (located in the ./scripts/ or ./materials/ directories) under the Apache License, Version 2.0.

Acknowledgments

This work was supported by a project from the Ministry of Education, Culture, Sports, Science, and Technology (MEXT) aiming at "establishment of research and development centers to ensure the transparency and reliability of generative AI models," along with other contributions.

We gratefully acknowledge Lianmin Zheng, the author of the original LMSYS-Chat-1M paper, for granting permission to distribute LMSYS-Chat-1M-Synth-Ja-and-En dataset as a derivative work of the original dataset.

End of document

Contributors

SM
Sakae Mizuki

42 commits

s-mizuki-nlp

15 commits

maym15

9 commits

YO
YoumiMa

8 commits

tokyotech-llm/lmsys-chat-1m-synth

Dataset

23

stars

76

commits

Feb 20, 2026

updated

README

LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M

This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M.

Additional Materials

We distribute Python scripts used to develop the dataset under the ./materials/ directory. The directory includes scripts for generating assistant responses and scoring preferences. These scripts are provided as-is solely for reproducibility of research purpose. We do not support or take responsibility in using these scripts.

License Information - Dataset

We publish the synthesized portion of the dataset under mixed licenses for each subset as follows:

User Instructions Translated into Japanese

The subsets lmsys-chat-1m-first-turn-user-instructions-ja.jsonl.gz.gpg and lmsys-chat-1m-first-turn-user-instructions-ja+unsafe.jsonl.gz.gpg, termed "Japanese Instructions," are distributed under the LMSYS-Chat-1M Dataset License Agreement. To access the original dataset and obtain the decryption key for the Japanese Instructions, you must agree to the license and provide your contact information. Please note that the "Right to Request Deletion" clause from the LMSYS-Chat-1M Dataset License also applies to Japanese Instructions: The original dataset authors retain the right to request you to delete all copies of the Japanese Instructions (in whole or in part) in your possession and control. You are required to comply with any and all such requests.

Assistant Responses and Preference Scores

The subset llama3.1-lmsys-chat-1m-synth-ja+en.jsonl.gz is distributed under the LLAMA 3.1 COMMUNITY LICENSE AGREEMENT.

The subsets gemma2-lmsys-chat-1m-synth-ja+en.jsonl.gz and gemma3-lmsys-chat-1m-synth-ja+en.jsonl.gz are distributed under the GEMMA TERMS OF USE.

The subset gpt-oss-lmsys-chat-1m-synth-ja+en.jsonl.gz is distributed under the Apache License, Version 2.0, in accordance with the gpt-oss license.

License Information - Scripts

We distribute the Python and Shell scripts (located in the ./scripts/ or ./materials/ directories) under the Apache License, Version 2.0.

Acknowledgments

This work was supported by a project from the Ministry of Education, Culture, Sports, Science, and Technology (MEXT) aiming at "establishment of research and development centers to ensure the transparency and reliability of generative AI models," along with other contributions.

We gratefully acknowledge Lianmin Zheng, the author of the original LMSYS-Chat-1M paper, for granting permission to distribute LMSYS-Chat-1M-Synth-Ja-and-En dataset as a derivative work of the original dataset.

End of document

Contributors

SM
Sakae Mizuki

42 commits

s-mizuki-nlp

15 commits

maym15

9 commits

YO
YoumiMa

8 commits