๐ค Baichuan-M1-14B-Base โข ๐ค Baichuan-M1-14B-Instruct โข ๐ Technical Report โข ๐ฌ WeChat
Baichuan-14B-M1 is the industry's first open-source large language model developed from scratch by Baichuan Intelligence, specifically optimized for medical scenarios. While excelling in general capabilities, it demonstrates powerful performance in the medical field. It achieves results comparable to models of similar size in most general benchmark evaluations, while outperforming models five times larger in medical scenarios. Below are the core features of the model:
We conducted meticulous data collection and synthesis for the medical field, including:
We innovatively adopted a multi-stage curriculum learning and alignment optimization approach, systematically enhancing model capabilities through the following two parts:
Training is divided into three stages, progressively optimizing the model's general and medical domain capabilities:
Enhancing model generation quality, logical reasoning, and user preference alignment through reinforcement learning and pairwise data optimization:
This combined approach of multi-stage and alignment optimization enables the model to achieve exceptional performance in both general and medical domain capabilities.
Our evaluation covers all mainstream benchmarks, achieving excellent metrics in both open-source and closed-source evaluations, demonstrating outstanding medical scenario capabilities while maintaining strong general performance.
| Category | Benchmark | Baichuan-M1-14B-Instruct | Qwen2.5-14B-Instruct | Qwen2.5-72B-Instruct | claude-3.5-sonnet-20241022 | gpt-4o |
|---|---|---|---|---|---|---|
| Average Score | 72.23 | 65.39 | 70.51 | 74.85 | 75.00 | |
| Clinical Practice | cmbclin | 77.40 | 71.51 | 75.36 | 78.37 | 75.36 |
| clinicalbench_diag | 70.90 | 68.85 | 72.23 | 75.00 | 73.05 | |
| clinicalbench_hos | 70.05 | 68.83 | 70.53 | 65.58 | 69.38 | |
| clinicalbench_treat | 56.38 | 55.03 | 57.30 | 64.03 | 59.35 | |
| rarearena_rdc | 81.80 | 66.40 | 76.20 | 89.60 | 88.40 | |
| rarearena_rds | 54.00 | 42.60 | 49.80 | 59.80 | 57.20 | |
| rarebench | 59.60 | 52.80 | 60.60 | 65.30 | 62.80 | |
| Exams | cmexam | 80.10 | 77.70 | 82.70 | 77.50 | 78.00 |
| Pediatric Qualification Exam | 78.48 | 74.68 | 84.81 | 76.58 | 78.48 | |
| Internal Medicine Qualification Exam | 83.42 | 86.10 | 87.17 | 87.70 | 83.42 | |
| General Practice Qualification Exam | 87.07 | 88.44 | 88.44 | 81.63 | 84.35 | |
| USMLE | 78.00 | 67.20 | 76.70 | 85.90 | 87.10 | |
| medbullets | 66.88 | 54.22 | 64.29 | 72.40 | 75.97 | |
| mediq | 83.40 | 66.80 | 79.90 | 88.80 | 90.20 | |
| nejmqa | 49.75 | 45.69 | 50.76 | 69.54 | 54.31 | |
| pubmedqa | 75.20 | 76.40 | 75.60 | 77.00 | 77.60 | |
| redisqa | 74.50 | 69.70 | 75.00 | 83.20 | 82.80 | |
| Basic Capabilities | mednli_dis | 80.40 | 68.90 | 74.90 | 58.30 | 79.80 |
| medcalc | 56.00 | 31.40 | 37.90 | 52.60 | 49.00 | |
| MMLU-anatomy | 80.00 | 67.41 | 71.11 | 86.67 | 91.11 | |
| MMLU-virology | 54.82 | 56.02 | 53.01 | 54.22 | 57.23 | |
| MMLU-genetics | 91.00 | 82.00 | 87.00 | 97.00 | 95.00 | |
We recommend using the latest version of the Transformers library (at least 4.47.0). The following code snippet demonstrates how to use the Baichuan-M1-14B-Instruct model:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# 1. Load pre-trained model and tokenizer
model_name = "baichuan-inc/Baichuan-M1-14B-Base"
tokenizer = AutoTokenizer.from_pretrained(model_name,trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_name,trust_remote_code=True,torch_dtype = torch.bfloat16).cuda()
input_text = "I have recently recovered from my cold."
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)
outputs = model.generate(
inputs["input_ids"],
max_length=100,
)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print("Generated Text:")
print(generated_text)
The use of the model must comply with ใBaichuan-M1-14Bๆจกๅ็คพๅบ่ฎธๅฏๅ่ฎฎใ.
The development team of Baichuan has not developed any commercial applications based on this model. All users must comply with laws and regulations and must not use the model for harmful national security or illegal purposes.
If you need to cite our work, please use the following reference:
@article{baichuan-m1-2025,
title={Baichuan-M1: Pushing the Medical Capability of Large Language Models},
author={Bingning Wang, Haizhou Zhao, Huozhi Zhou, Liang Song, Mingyu Xu, Wei Cheng, Xiangrong Zeng, Yupeng Zhang, Yuqi Huo, Zecheng Wang, Zhengyun Zhao and others},
journal={arXiv preprint arXiv:2502.12671},
year={2025}
}
๐ค Baichuan-M1-14B-Base โข ๐ค Baichuan-M1-14B-Instruct โข ๐ Technical Report โข ๐ฌ WeChat
Baichuan-14B-M1 is the industry's first open-source large language model developed from scratch by Baichuan Intelligence, specifically optimized for medical scenarios. While excelling in general capabilities, it demonstrates powerful performance in the medical field. It achieves results comparable to models of similar size in most general benchmark evaluations, while outperforming models five times larger in medical scenarios. Below are the core features of the model:
We conducted meticulous data collection and synthesis for the medical field, including:
We innovatively adopted a multi-stage curriculum learning and alignment optimization approach, systematically enhancing model capabilities through the following two parts:
Training is divided into three stages, progressively optimizing the model's general and medical domain capabilities:
Enhancing model generation quality, logical reasoning, and user preference alignment through reinforcement learning and pairwise data optimization:
This combined approach of multi-stage and alignment optimization enables the model to achieve exceptional performance in both general and medical domain capabilities.
Our evaluation covers all mainstream benchmarks, achieving excellent metrics in both open-source and closed-source evaluations, demonstrating outstanding medical scenario capabilities while maintaining strong general performance.
| Category | Benchmark | Baichuan-M1-14B-Instruct | Qwen2.5-14B-Instruct | Qwen2.5-72B-Instruct | claude-3.5-sonnet-20241022 | gpt-4o |
|---|---|---|---|---|---|---|
| Average Score | 72.23 | 65.39 | 70.51 | 74.85 | 75.00 | |
| Clinical Practice | cmbclin | 77.40 | 71.51 | 75.36 | 78.37 | 75.36 |
| clinicalbench_diag | 70.90 | 68.85 | 72.23 | 75.00 | 73.05 | |
| clinicalbench_hos | 70.05 | 68.83 | 70.53 | 65.58 | 69.38 | |
| clinicalbench_treat | 56.38 | 55.03 | 57.30 | 64.03 | 59.35 | |
| rarearena_rdc | 81.80 | 66.40 | 76.20 | 89.60 | 88.40 | |
| rarearena_rds | 54.00 | 42.60 | 49.80 | 59.80 | 57.20 | |
| rarebench | 59.60 | 52.80 | 60.60 | 65.30 | 62.80 | |
| Exams | cmexam | 80.10 | 77.70 | 82.70 | 77.50 | 78.00 |
| Pediatric Qualification Exam | 78.48 | 74.68 | 84.81 | 76.58 | 78.48 | |
| Internal Medicine Qualification Exam | 83.42 | 86.10 | 87.17 | 87.70 | 83.42 | |
| General Practice Qualification Exam | 87.07 | 88.44 | 88.44 | 81.63 | 84.35 | |
| USMLE | 78.00 | 67.20 | 76.70 | 85.90 | 87.10 | |
| medbullets | 66.88 | 54.22 | 64.29 | 72.40 | 75.97 | |
| mediq | 83.40 | 66.80 | 79.90 | 88.80 | 90.20 | |
| nejmqa | 49.75 | 45.69 | 50.76 | 69.54 | 54.31 | |
| pubmedqa | 75.20 | 76.40 | 75.60 | 77.00 | 77.60 | |
| redisqa | 74.50 | 69.70 | 75.00 | 83.20 | 82.80 | |
| Basic Capabilities | mednli_dis | 80.40 | 68.90 | 74.90 | 58.30 | 79.80 |
| medcalc | 56.00 | 31.40 | 37.90 | 52.60 | 49.00 | |
| MMLU-anatomy | 80.00 | 67.41 | 71.11 | 86.67 | 91.11 | |
| MMLU-virology | 54.82 | 56.02 | 53.01 | 54.22 | 57.23 | |
| MMLU-genetics | 91.00 | 82.00 | 87.00 | 97.00 | 95.00 | |
We recommend using the latest version of the Transformers library (at least 4.47.0). The following code snippet demonstrates how to use the Baichuan-M1-14B-Instruct model:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# 1. Load pre-trained model and tokenizer
model_name = "baichuan-inc/Baichuan-M1-14B-Base"
tokenizer = AutoTokenizer.from_pretrained(model_name,trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_name,trust_remote_code=True,torch_dtype = torch.bfloat16).cuda()
input_text = "I have recently recovered from my cold."
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)
outputs = model.generate(
inputs["input_ids"],
max_length=100,
)
generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
print("Generated Text:")
print(generated_text)
The use of the model must comply with ใBaichuan-M1-14Bๆจกๅ็คพๅบ่ฎธๅฏๅ่ฎฎใ.
The development team of Baichuan has not developed any commercial applications based on this model. All users must comply with laws and regulations and must not use the model for harmful national security or illegal purposes.
If you need to cite our work, please use the following reference:
@article{baichuan-m1-2025,
title={Baichuan-M1: Pushing the Medical Capability of Large Language Models},
author={Bingning Wang, Haizhou Zhao, Huozhi Zhou, Liang Song, Mingyu Xu, Wei Cheng, Xiangrong Zeng, Yupeng Zhang, Yuqi Huo, Zecheng Wang, Zhengyun Zhao and others},
journal={arXiv preprint arXiv:2502.12671},
year={2025}
}