hongzhouyu/FineMed-DPO

Dataset

Introduction

2

16 commits

1 linked in READMEs

updated Feb 12, 2025

See the code

README

Introduction

This dataset is constructed using Qwen2.5-72B-Instruct and QwQ-32B-Preview, and it serves as the foundation for fine-tuning FineMedLM-o1.

This repository contains the DPO data used during the training process. To access the dataset, you can either use the load_dataset function or directly download the parquet files from the designated folder.

For details, see our paper and GitHub repository.

Citation

If you find our data useful, please consider citing our work!

@misc{yu2025finemedlmo1enhancingmedicalreasoning,
    title={FineMedLM-o1: Enhancing the Medical Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training}, 
    author={Hongzhou Yu and Tianhao Cheng and Ying Cheng and Rui Feng},
    year={2025},
    eprint={2501.09213},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2501.09213}, 
}
medical

hongzhouyu/FineMed-DPO

Dataset

Introduction

2

16 commits

1 linked in READMEs

updated Feb 12, 2025

See the code

README

Introduction

This dataset is constructed using Qwen2.5-72B-Instruct and QwQ-32B-Preview, and it serves as the foundation for fine-tuning FineMedLM-o1.

This repository contains the DPO data used during the training process. To access the dataset, you can either use the load_dataset function or directly download the parquet files from the designated folder.

For details, see our paper and GitHub repository.

Citation

If you find our data useful, please consider citing our work!

@misc{yu2025finemedlmo1enhancingmedicalreasoning,
    title={FineMedLM-o1: Enhancing the Medical Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training}, 
    author={Hongzhou Yu and Tianhao Cheng and Ying Cheng and Rui Feng},
    year={2025},
    eprint={2501.09213},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2501.09213}, 
}
medical