datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
symptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/symptom_to_diagnosis.menopause-symptoms
Perimenopause & Menopause Symptoms — Science-Cited Reference Dataset
An openly-licensed, structured dataset of 193 perimenopause and menopause symptom explainers, each mapped to its source article and to peer-reviewed citations (DOIs). Built for researchers, developers, and anyone building menopause health tools or AI assistants who want a clean, attributable symptom reference.
What's inside
193 rows (one per symptom / topic explainer)
193 rows carry one or more… See the full description on the dataset page: https://huggingface.co/datasets/Whiterocket/menopause-symptoms.symptom-based-disease-prediction-v1
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1
📘 Overview
This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs).
Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v1.symptoms_disease_v1SymptomsDisease246k
Source
Disease-Symptom-Extensive-Clean
Context Sample
{
"query": "Having these specific symptoms: anxiety and nervousness, depression, shortness of breath, depressive or psychotic symptoms, dizziness, palpitations, irregular heartbeat, breathing fast may indicate",
"response": "You may have panic disorder"
}
Raw Sample
{
"query": "dizziness, abnormal involuntary movements, headache, diminished vision",
"response": "pseudotumor cerebri"
}
symptom-based-disease-prediction-v2
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v2
📘 Overview
Version 2 of the Symptom-Based Disease Prediction Dataset is an improved and more structured medical reasoning dataset designed for Large Language Models (LLMs) and AI-driven healthcare research.
This version introduces:
cleaner formatting,
improved disease grouping,
better confidence separation,
enhanced consistency,
and more fine-tuning-friendly outputs.
Each sample presents symptoms in… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v2.datatager_symptom_recognition_and_advice
If you like our project, please give us a star ⭐
[GitHub | DataTager Home]
Extract Medical Information Dataset
Prompt for Training
When training your model with this dataset, prepend the following prompt to each input instance:
你需要去做的是理解患者的咨询文本,并基于这些症状提供一个可能的医学解释以及相应的建议措施。请始终确保你的输出中包括以下元素:1. 对输入中提到的症状的识别和确认。2. 基于症状的可能医学解释。3. 针对进一步诊断或治疗的建议措施。
Description
AnyTaskTune is a publication by the DataTager team. We advocate for rapid training of large… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/datatager_symptom_recognition_and_advice.MedQA_SymptomDisease_small_DutchA Dutch translation of this huggingface dataset using GPT4.1 mini, with the courtesy of Prognosis.
Symptoms_to_disease_7ksymptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/aryang2005/symptom_to_diagnosis.symptom-diagnoser-sft
Symptom to Diagnosis (ChatDoctor-200k)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Patient symptom descriptions → differential diagnosis with ICD-10 codes
Why download this
Build clinical decision support chatbots, symptom checkers, or medical question-answering systems. Largest dataset in the AxisMapper suite.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/symptom-diagnoser-sft.symptom-based-disease-prediction-v2
Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v1
📘 Overview
This dataset is an early-stage (Version 1) medical dataset created for symptom-based disease prediction using Large Language Models (LLMs).
Each record presents patient symptoms in an instruction-style prompt and returns multiple possible diseases grouped by confidence levels.The primary goal of this version is to establish structure, consistency, and reasoning format, not final model… See the full description on the dataset page: https://huggingface.co/datasets/RonalLI/symptom-based-disease-prediction-v2.adaption-vaidya-rural-symptoms
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vaidya_rural_symptoms
This dataset maps colloquial symptom expressions from multiple Indian languages and dialects to standardized medical meanings and severity levels. It is designed to bridge the gap between rural healthcare communication and formal medical terminology for NLP applications. The data includes core fields for symptom phrases, language, dialect, corrected meaning, and… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-vaidya-rural-symptoms.json_symptomsymptoms-disease_dataset_for_LLMadaption-indian-symptom-phrases
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-indian_symptom_phrases
This dataset contains pairs of colloquial symptom descriptions and their standardized medical equivalents in various Indian languages, including Hindi, Marathi, Malayalam, Telugu, Kannada, and Gujarati. Each entry maps a common user phrase describing a physical ailment to a formal term with an English translation. The collection is designed to support natural… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-indian-symptom-phrases.symptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/Bhupesh18/symptom_to_diagnosis.adaption-vaidya-rural-symptoms-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vaidya_rural_symptoms
This dataset maps colloquial symptom expressions from multiple Indian languages and dialects to standardized medical meanings and severity levels. It is designed to bridge the gap between rural healthcare communication and formal medical terminology for NLP applications. The data includes core fields for symptom phrases, language, dialect, corrected meaning, and… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-vaidya-rural-symptoms-v1.symptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/HNJ1998/symptom_to_diagnosis.symptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/Shashwat250/symptom_to_diagnosis.symptom_to_diagnosis
Dataset Summary
This dataset contains natural language descriptions of symptoms labeled with 22 corresponding diagnoses. Gretel/symptom_to_diagnosis provides 1065 symptom descriptions in the English language labeled with 22 diagnoses, focusing on fine-grained single-domain diagnosis.
Data Fields
Each row contains the following fields:
input_text : A string field containing symptoms
output_text : A string field containing a diagnosis
Example:
{
"output_text": "drug… See the full description on the dataset page: https://huggingface.co/datasets/john-hemedy/symptom_to_diagnosis.JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTEDThis dataset is high-consistency instruction tuning dataset.
converting messy, subjective human health-style text → structured, non-diagnostic extraction format
1.Literal extraction discipline
2.Source separation logic - very strong schema grounding training if DPO
3.Anti-inference constraint
1.High ambiguity coverage
2.Contradiction handling included
3.Minimization bias detection
This dataset is a STRICT schema regulation.
Does well at:
strict extraction
preserving uncertainty words… See the full description on the dataset page: https://huggingface.co/datasets/sadnjasdkn/JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTED.preprocessed_json_patients_symptoms_to_diagnosis_symptoms_datasetrpancreatic_cancer_symptomssymptomsheart-arrhythmias-symptomssymptoms
