๐ Synthetic Multipersona Doctor Patient Conversations.
35
11 commits
1 linked in READMEs
updated Jan 29, 2025
Author: Nisten Tahiraj
License: MIT
๐ง Conversations generated in the Following languages
_
English
Chinese
Japanese
Danish
German
French
_
More languages coming :) Follow our org lead by Doctor @JohnsonThomasMD for more updates, DeepSeek R1 generations and a new mobile opensource medical model are in the works too ๐ .
Before the data was generated the medical performance of the LLM was measured to be significantly higher than even Google's MedPalm 2.
Reference: MedPalm two scores no higher than 72% https://paperswithcode.com/sota/multiple-choice-question-answering-mcqa-on-21
Despite the driver issues, deepseek v3 instruct has stellar scores in medical benmarking, here running in fp8_w8a8 on 8x AMD Mi300x card the multimedqa bench. Little to no difference was observed in medical benchmarking in bfloat16 vs 8bit. However other tests showed some divergence: https://x.com/nisten/status/1874996106540503367
Yes, raw deepseek v3 with no special prompting scores 79% vs only 72% for the complicated CoT MedPalm2 API setup.
The newer DeepSeek R1 has not yet been tested.
Feel free to leave comments, concerns, and even contribute more data to open science.

11 commits
๐ Synthetic Multipersona Doctor Patient Conversations.
35
11 commits
1 linked in READMEs
updated Jan 29, 2025
Author: Nisten Tahiraj
License: MIT
๐ง Conversations generated in the Following languages
_
English
Chinese
Japanese
Danish
German
French
_
More languages coming :) Follow our org lead by Doctor @JohnsonThomasMD for more updates, DeepSeek R1 generations and a new mobile opensource medical model are in the works too ๐ .
Before the data was generated the medical performance of the LLM was measured to be significantly higher than even Google's MedPalm 2.
Reference: MedPalm two scores no higher than 72% https://paperswithcode.com/sota/multiple-choice-question-answering-mcqa-on-21
Despite the driver issues, deepseek v3 instruct has stellar scores in medical benmarking, here running in fp8_w8a8 on 8x AMD Mi300x card the multimedqa bench. Little to no difference was observed in medical benchmarking in bfloat16 vs 8bit. However other tests showed some divergence: https://x.com/nisten/status/1874996106540503367
Yes, raw deepseek v3 with no special prompting scores 79% vs only 72% for the complicated CoT MedPalm2 API setup.
The newer DeepSeek R1 has not yet been tested.
Feel free to leave comments, concerns, and even contribute more data to open science.

11 commits