FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT

Dataset

13

stars

5

commits

2

linked in READMEs

Aug 25, 2025

updated

Multimodal Data
Traditional Chinese Medicine

README

πŸ“š Introduction

This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.

For details, see our paper and GitHub repository.

πŸ“Š Dataset Overview

The open-sourced fine-tuning dataset consists of three parts:

ModalityData Quantity
TCM Text InstructionsπŸ“ Text87K
TCM Visual InstructionsπŸ“ Text, πŸ‘οΈ Visual67K
TCM Speech InstructionsπŸ“ Text, πŸ‘οΈ Visual, πŸŽ™οΈ Audio91K

⚠️ Note: Since TCM signal datasets, such as pulse and smell, involve private information, we recommend users download them from the corresponding paper.

πŸ“– Citation

If you find our data useful, please consider citing our work!

@misc{chen2025shizhengptmultimodalllmstraditional,
      title={ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine}, 
      author={Junying Chen and Zhenyang Cai and Zhiheng Liu and Yunjin Yang and Rongsheng Wang and Qingying Xiao and Xiangyi Feng and Zhan Su and Jing Guo and Xiang Wan and Guangjun Yu and Haizhou Li and Benyou Wang},
      year={2025},
      eprint={2508.14706},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2508.14706},
}

Contributors

jymcc

5 commits

FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT

Dataset

13

stars

5

commits

2

linked in READMEs

Aug 25, 2025

updated

Multimodal Data
Traditional Chinese Medicine

README

πŸ“š Introduction

This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.

For details, see our paper and GitHub repository.

πŸ“Š Dataset Overview

The open-sourced fine-tuning dataset consists of three parts:

ModalityData Quantity
TCM Text InstructionsπŸ“ Text87K
TCM Visual InstructionsπŸ“ Text, πŸ‘οΈ Visual67K
TCM Speech InstructionsπŸ“ Text, πŸ‘οΈ Visual, πŸŽ™οΈ Audio91K

⚠️ Note: Since TCM signal datasets, such as pulse and smell, involve private information, we recommend users download them from the corresponding paper.

πŸ“– Citation

If you find our data useful, please consider citing our work!

@misc{chen2025shizhengptmultimodalllmstraditional,
      title={ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine}, 
      author={Junying Chen and Zhenyang Cai and Zhiheng Liu and Yunjin Yang and Rongsheng Wang and Qingying Xiao and Xiangyi Feng and Zhan Su and Jing Guo and Xiang Wan and Guangjun Yu and Haizhou Li and Benyou Wang},
      year={2025},
      eprint={2508.14706},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2508.14706},
}

Contributors

jymcc

5 commits