Coldog2333/JMedBench

Dataset

7

stars

139

commits

5

linked in READMEs

Jan 19, 2025

updated

README

Maintainers

  • Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp
  • Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp

If you find any error in this benchmark or want to contribute to this benchmark, please feel free to contact us.

Introduction

This is a dataset collection of JMedBench, which is a benchmark for evaluating Japanese biomedical large language models (LLMs). Details can be found in this paper. We also provide an evaluation framework, med-eval, for easy evaluation.

The JMedBench consists of 20 datasets across 5 tasks, listed below.

TaskDatasetLicenseSourceNote
MCQAmedmcqa_jpMITMedMCQATranslated
usmleqa_jpMITMedQATranslated
medqa_jpMITMedQATranslated
mmlu_medical_jpMITMMLUTranslated
jmmlu_medicalCC-BY-SA-4.0JMMLU
igakuqa-paper
pubmedqa_jpMITPubMedQATranslated
MTejmmtCC-BY-4.0paper
NERmrner_medicineCC-BY-4.0JMED-LLM
mrner_diseaseCC-BY-4.0JMED-LLM
nrnerCC-BY-NC-SA-4.0JMED-LLM
bc2gm_jpUnknownBLURBTranslated
bc5chem_jpOtherBLURBTranslated
bc5disease_jpOtherBLURBTranslated
jnlpba_jpUnknownBLURBTranslated
ncbi_disease_jpUnknownBLURBTranslated
DCcradeCC-BY-4.0JMED-LLM
rrtnmCC-BY-4.0JMED-LLM
smdisCC-BY-4.0JMED-LLM
STSjcstsCC-BY-NC-SA-4.0paper

Limitations

Please be aware of the risks, biases, and limitations of this benchmark.

As introduced in the previous section, some evaluation datasets are translated from the original sources (in English). Although we used the most powerful API from OpenAI (i.e., gpt-4-0613) to conduct the translation, it may be unavoidable to contain incorrect or inappropriate translations. If you are developing biomedical LLMs for real-world applications, please conduct comprehensive human evaluation before deployment.

Citation

BibTeX: If our JMedBench is helpful for you, please cite our work:

@misc{jiang2024jmedbenchbenchmarkevaluatingjapanese,
      title={JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models}, 
      author={Junfeng Jiang and Jiahao Huang and Akiko Aizawa},
      year={2024},
      eprint={2409.13317},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2409.13317}, 
}

Contributors

CO
Coldog2333

139 commits

Coldog2333/JMedBench

Dataset

7

stars

139

commits

5

linked in READMEs

Jan 19, 2025

updated

README

Maintainers

  • Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp
  • Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp

If you find any error in this benchmark or want to contribute to this benchmark, please feel free to contact us.

Introduction

This is a dataset collection of JMedBench, which is a benchmark for evaluating Japanese biomedical large language models (LLMs). Details can be found in this paper. We also provide an evaluation framework, med-eval, for easy evaluation.

The JMedBench consists of 20 datasets across 5 tasks, listed below.

TaskDatasetLicenseSourceNote
MCQAmedmcqa_jpMITMedMCQATranslated
usmleqa_jpMITMedQATranslated
medqa_jpMITMedQATranslated
mmlu_medical_jpMITMMLUTranslated
jmmlu_medicalCC-BY-SA-4.0JMMLU
igakuqa-paper
pubmedqa_jpMITPubMedQATranslated
MTejmmtCC-BY-4.0paper
NERmrner_medicineCC-BY-4.0JMED-LLM
mrner_diseaseCC-BY-4.0JMED-LLM
nrnerCC-BY-NC-SA-4.0JMED-LLM
bc2gm_jpUnknownBLURBTranslated
bc5chem_jpOtherBLURBTranslated
bc5disease_jpOtherBLURBTranslated
jnlpba_jpUnknownBLURBTranslated
ncbi_disease_jpUnknownBLURBTranslated
DCcradeCC-BY-4.0JMED-LLM
rrtnmCC-BY-4.0JMED-LLM
smdisCC-BY-4.0JMED-LLM
STSjcstsCC-BY-NC-SA-4.0paper

Limitations

Please be aware of the risks, biases, and limitations of this benchmark.

As introduced in the previous section, some evaluation datasets are translated from the original sources (in English). Although we used the most powerful API from OpenAI (i.e., gpt-4-0613) to conduct the translation, it may be unavoidable to contain incorrect or inappropriate translations. If you are developing biomedical LLMs for real-world applications, please conduct comprehensive human evaluation before deployment.

Citation

BibTeX: If our JMedBench is helpful for you, please cite our work:

@misc{jiang2024jmedbenchbenchmarkevaluatingjapanese,
      title={JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models}, 
      author={Junfeng Jiang and Jiahao Huang and Akiko Aizawa},
      year={2024},
      eprint={2409.13317},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2409.13317}, 
}

Contributors

CO
Coldog2333

139 commits