CoolFace
Datasetpublic

tomshe/Pediatrics_questions

Dataset Card for Pediatrics MCQ Dataset Details Dataset Description This dataset comprises high-quality multiple-choice questions (MCQs) covering core biomedical knowledge and clinical scenarios from pediatrics. It includes 50 questions, each with four possible answer choices. These questions were specifically curated for research evaluating pediatric medical knowledge, clinical reasoning, and confidence-based interactions among medical trainees and… See the full description on the dataset page: https://huggingface.co/datasets/tomshe/Pediatrics_questions.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes11downloads
Dataset Card

Dataset Card for Pediatrics MCQ

Dataset Details

Dataset Description

This dataset comprises high-quality multiple-choice questions (MCQs) covering core biomedical knowledge and clinical scenarios from pediatrics. It includes 50 questions, each with four possible answer choices. These questions were specifically curated for research evaluating pediatric medical knowledge, clinical reasoning, and confidence-based interactions among medical trainees and large language models (LLMs).


Uses

Direct Use

This dataset is suitable for:

  • Evaluating pediatric medical knowledge and clinical reasoning skills of medical students and healthcare professionals.
  • Benchmarking performance and reasoning capabilities of large language models (LLMs) in pediatric medical question-answering tasks.
  • Research on collaborative human–AI and human–human interactions involving pediatric clinical decision-making.

Out-of-Scope Use

  • Not intended as a diagnostic or clinical decision-making tool in real clinical settings.
  • Should not be used to train systems intended for direct clinical application without extensive validation.

Dataset Structure

The dataset comprises 50 multiple-choice questions with four answer choices. The dataset includes the following fields:

  • question_text: The clinical vignette or biomedical question.
  • optionA: First possible answer choice.
  • optionB: Second possible answer choice.
  • optionC: Third possible answer choice.
  • optionD: Fourth possible answer choice.
  • answer: The correct answer text.
  • answer_idx: The correct answer choice (A, B, C, or D).

Dataset Creation

Curation Rationale

The dataset was created to study knowledge diversity, internal confidence, and collaborative decision-making between medical trainees and AI agents. Questions were carefully selected to represent authentic licensing exam–style questions in pediatrics, ensuring ecological validity for medical education and AI–human collaborative studies.


Source Data

Data Collection and Processing

The questions were sourced and adapted from standardized pediatric medical licensing preparation materials. All questions were reviewed, translated, and validated by licensed pediatricians.

Who are the source data producers?

The original data sources are standard pediatric medical licensing examination preparation materials.


Personal and Sensitive Information

The dataset does not contain any personal, sensitive, or identifiable patient or clinician information. All clinical scenarios are fictionalized or generalized for educational and research purposes.


Bias, Risks, and Limitations

  • The dataset size (50 questions) is limited; therefore, findings using this dataset might not generalize broadly.
  • Content is limited to pediatrics; results may not generalize across all medical specialties.

Citation

If using this dataset, please cite:


More Information

For more details, please contact the dataset author listed below.


Dataset Card Author

  • Tom Sheffer (The Hebrew University of Jerusalem)

Dataset Card Contact