CoolFace
Datasetpublic

Sabr-Research/SR-medical-sft-preview-reviewed

Overview This is a Supervised Fine-Tuning preview dataset consisting of structured clinical and forensic medical reasoning. It is designed to teach language models to analyze complex patient histories, physical presentations, and diagnostic findings using systematic differential analysis and pathophysiological synthesis within the context of peer-reviewed medical literature and case reports. The dataset is derived from real, open-access PubMed Central (PMC) medical case reports.… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-medical-sft-preview-reviewed.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes21downloads
Dataset Card

Overview

This is a Supervised Fine-Tuning preview dataset consisting of structured clinical and forensic medical reasoning. It is designed to teach language models to analyze complex patient histories, physical presentations, and diagnostic findings using systematic differential analysis and pathophysiological synthesis within the context of peer-reviewed medical literature and case reports.

The dataset is derived from real, open-access PubMed Central (PMC) medical case reports. Each row provides a concrete, grounded medical background prompt and a structured, zero-hallucination Chain-of-Thought (CoT) response that tracks the diagnostic journey from initial presentation to final diagnosis and management.

Size and Intended Use

This is a preview-scale dataset. It is intended for researchers and developers to explore this clinical use case, evaluate quality, test formatting, and prototype small-scale LoRA fine-tunes for medical reasoning, clinical text analysis, and medical forensic applications. It is not sized for training a production-ready model from scratch.

Dataset Structure

Each row in the .jsonl file contains:

  • —`case_id`: A unique identifier.
  • —`prompt`: A structured medical vignette consisting of explicit Background Facts, a Baseline Event, and a comprehensive Task paragraph summarizing clinical exam details, laboratory workups, imaging, and histopathological findings.
  • —`response`: A structured clinical Chain-of-Thought (CoT) analysis containing exactly four bullet points:
  • —Patient Context & Presentation: A summary of the patient's demographics, clinical timeline (e.g., acute vs. subacute), physical exam findings, and immediate clinical concerns.
  • —Differential Diagnosis (Eliminated Hypotheses): An explanation of secondary conditions, malignancies, or alternative infectious etiologies that were explicitly ruled out based on diagnostic, laboratory, or histopathological evidence.
  • —Diagnostic Sequence & Pathophysiology: A chronological synthesis of the clinical workup, detailing how specific diagnostic markers, imaging patterns, and microscopic examinations logically rule out alternatives and confirm the precise physiological or compressive mechanisms.
  • —Final Diagnosis & Outcome: The definitive final diagnosis followed by the specific therapeutic intervention, medication regimen, or forensic summary, including long-term patient outcomes where applicable.

Methodology & Provenance

Please note: The analyses in this dataset are not verbatim excerpts from medical journals. The data was generated by extracting core clinical facts, timeline events, and biological mechanisms using an automated data pipeline.

Verification: Each sample has been verified by human experts against the original open-access PMC case reports to ensure strict clinical accuracy and grounding.

Licensing Information and Credits

  • —Source Material Information and Credits: This dataset was produced by processing original medical case reports sourced from the PubMed Central Open Access Subset. The source texts utilized to generate this dataset were strictly filtered to include only those distributed under Creative Commons Zero (CC0) and Creative Commons Attribution (CC-BY) licenses. To comply with the attribution requirements of the CC-BY license, a supplementary mapping file (credit.csv) is provided alongside this dataset, which links every case_id directly to its original publication URL, allowing users to identify and credit the original authors and copyright holders.
  • —Dataset Generations: The dataset is released openly under the Apache 2.0 License.

Commercial Scaling & Contact

This dataset was produced by Sabr Research. We build reasoning datasets verified by human experts to train AI models at scale.

If you require scaled versions of this dataset (thousands of rows), or wish to engage in a customized partnership to build specialized datasets, please reach out to us at Sabr Research.