*Note: these datasets were uploaded on March 19th, 2024 by Hugging Face staff from the ancillary files attached to the original arXiv submission.*
7
10 commits
1 linked in READMEs
updated Mar 19, 2024
Note: these datasets were uploaded on March 19th, 2024 by Hugging Face staff from the ancillary files attached to the original arXiv submission.
https://arxiv.org/abs/2403.12025
We include adversarial questions for each of the seven EquityMedQA datasets: OMAQ, EHAI, FBRT-Manual, FBRT-LLM, TRINDS, CC-Manual, and CC-LLM. For FBRT-LLM, we include both the full set and the subset we sampled for evaluation in the work. For CC-Manual and CC-LLM, we provide two related questions on each line in their respective files. Data generated as a part of the empirical study (Med-PaLM 2 model outputs and human ratings) are not included in EquityMedQA.
We also include other datasets evaluated in this work: MultiMedQA, Mixed MMQA-OMAQ, and Omiye et al. These datasets are derived from:
See the paper for details on all datasets.
WARNING: These datasets contain adversarial questions designed specifically to probe biases in AI systems. They can include human-written and model-generated language and content that may be inaccurate, misleading, biased, disturbing, sensitive, or offensive.
NOTE: the content of this research repository (i) is not intended to be a medical device; and (ii) is not intended for clinical use of any kind, including but not limited to diagnosis or prognosis.
10 commits
*Note: these datasets were uploaded on March 19th, 2024 by Hugging Face staff from the ancillary files attached to the original arXiv submission.*
7
10 commits
1 linked in READMEs
updated Mar 19, 2024
Note: these datasets were uploaded on March 19th, 2024 by Hugging Face staff from the ancillary files attached to the original arXiv submission.
https://arxiv.org/abs/2403.12025
We include adversarial questions for each of the seven EquityMedQA datasets: OMAQ, EHAI, FBRT-Manual, FBRT-LLM, TRINDS, CC-Manual, and CC-LLM. For FBRT-LLM, we include both the full set and the subset we sampled for evaluation in the work. For CC-Manual and CC-LLM, we provide two related questions on each line in their respective files. Data generated as a part of the empirical study (Med-PaLM 2 model outputs and human ratings) are not included in EquityMedQA.
We also include other datasets evaluated in this work: MultiMedQA, Mixed MMQA-OMAQ, and Omiye et al. These datasets are derived from:
See the paper for details on all datasets.
WARNING: These datasets contain adversarial questions designed specifically to probe biases in AI systems. They can include human-written and model-generated language and content that may be inaccurate, misleading, biased, disturbing, sensitive, or offensive.
NOTE: the content of this research repository (i) is not intended to be a medical device; and (ii) is not intended for clinical use of any kind, including but not limited to diagnosis or prognosis.
10 commits