YuexingHao/MedPAIR
Dataset Card for MedPAIR MedPAIR represents a "Medical Dataset Comparing Physician Trainees and AI Relevance Estimation and Question Answering". We design MedPAIR to compare LLM reasoning processes to those of physician trainees and to enable future research to focus on relevant features. MedPAIR is a first benchmark step to matching the relevancy annotated by clinical professional labelers to that estimated by LLMs. The motivation for MedPAIR is to ensure that what the LLM… See the full description on the dataset page: https://huggingface.co/datasets/YuexingHao/MedPAIR.
Dataset Card for MedPAIR
MedPAIR represents a "Medical Dataset Comparing Physician Trainees and AI Relevance Estimation and Question Answering". We design MedPAIR to compare LLM reasoning processes to those of physician trainees and to enable future research to focus on relevant features. MedPAIR is a first benchmark step to matching the relevancy annotated by clinical professional labelers to that estimated by LLMs. The motivation for MedPAIR is to ensure that what the LLM finds relevant in a clinical case closely matches what a physician trainee finds relevant.
Dataset Details
There are two sequential stages: first, physician trainee labelers answered the QA question in column "q1"; then labeled each sentence's relevancy level in "Patient Profile".
We recruited 36 labelers (medical students or above) to curate the dataset. After filtering incorrect answers and their sentence labels, we got 1404 unique QAs with more than one correct labels. Among them, 104 QAs have only "High Relevance" in the "Patient Profile" and didn't show removing the "Low Relevance" and "Irrelevant" sentences could help them with decision-making. Therefore, we curate the total of 1,300 unique QA pairs with at least one "low relevance" or "irrelevant" sentence removed.
Dataset Sources
- Repository: https://github.com/YuexingHao/MedPAIR
- Paper: TBD
- Website: https://medpair.csail.mit.edu/
- Curated by: This project was primarily conducted and recieved ethics exempt via MIT COUHES. The project was assisted by researchers at other various academic and industry institutions.
- Funded by: This work was supported in part by an award from the Hasso Plattner Foundation, a National Science Foundation (NSF) CAREER Award (#2339381), and an AI2050 Early Career Fellowship (G-25-68042).
- Language(s): The dataset is in English.
- License: Physician trainees' labels within the dataset are licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
Terms of Use
Purpose
The Dataset is provided for the purpose of research and educational use in the field of input context relevance estimates, medical question answering, social science and related areas; and can be used to develop or evaluate artificial intelligence, including Large Language Models (LLMs).
Usage Restrictions
Users of the Dataset should adhere to the terms of use for a specific model when using its generated responses. This includes respecting any limitations or use case prohibitions set forth by the original model's creators or licensors.
Content Warning
The Dataset contains raw conversations that may include content considered unsafe or offensive. Users must apply appropriate filtering and moderation measures when using this Dataset for training purposes to ensure the generated outputs align with ethical and safety standards.
No Endorsement of Content
The conversations and data within this Dataset do not reflect the views or opinions of the Dataset creators, funders or any affiliated institutions. The dataset is provided as a neutral resource for research and should not be construed as endorsing any specific viewpoints.
Limitation of Liability
The authors and funders of this Dataset will not be liable for any claims, damages, or other liabilities arising from the use of the dataset, including but not limited to the misuse, interpretation, or reliance on any data contained within.
Issue Reporting
If there are any issues with the dataset, please report it to us via email yuexing@mit.edu.
