datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
menopause-symptoms
Perimenopause & Menopause Symptoms — Science-Cited Reference Dataset
An openly-licensed, structured dataset of 193 perimenopause and menopause symptom explainers, each mapped to its source article and to peer-reviewed citations (DOIs). Built for researchers, developers, and anyone building menopause health tools or AI assistants who want a clean, attributable symptom reference.
What's inside
193 rows (one per symptom / topic explainer)
193 rows carry one or more… See the full description on the dataset page: https://huggingface.co/datasets/Whiterocket/menopause-symptoms.symptoms_disease_v1SymptomsDisease246k
Source
Disease-Symptom-Extensive-Clean
Context Sample
{
"query": "Having these specific symptoms: anxiety and nervousness, depression, shortness of breath, depressive or psychotic symptoms, dizziness, palpitations, irregular heartbeat, breathing fast may indicate",
"response": "You may have panic disorder"
}
Raw Sample
{
"query": "dizziness, abnormal involuntary movements, headache, diminished vision",
"response": "pseudotumor cerebri"
}
Symptoms_to_disease_7kadaption-vaidya-rural-symptoms
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vaidya_rural_symptoms
This dataset maps colloquial symptom expressions from multiple Indian languages and dialects to standardized medical meanings and severity levels. It is designed to bridge the gap between rural healthcare communication and formal medical terminology for NLP applications. The data includes core fields for symptom phrases, language, dialect, corrected meaning, and… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-vaidya-rural-symptoms.symptoms-disease_dataset_for_LLMadaption-vaidya-rural-symptoms-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-vaidya_rural_symptoms
This dataset maps colloquial symptom expressions from multiple Indian languages and dialects to standardized medical meanings and severity levels. It is designed to bridge the gap between rural healthcare communication and formal medical terminology for NLP applications. The data includes core fields for symptom phrases, language, dialect, corrected meaning, and… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-vaidya-rural-symptoms-v1.JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTEDThis dataset is high-consistency instruction tuning dataset.
converting messy, subjective human health-style text → structured, non-diagnostic extraction format
1.Literal extraction discipline
2.Source separation logic - very strong schema grounding training if DPO
3.Anti-inference constraint
1.High ambiguity coverage
2.Contradiction handling included
3.Minimization bias detection
This dataset is a STRICT schema regulation.
Does well at:
strict extraction
preserving uncertainty words… See the full description on the dataset page: https://huggingface.co/datasets/sadnjasdkn/JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTED._symptoms_datasetpreprocessed_json_patients_symptoms_to_diagnosisrpancreatic_cancer_symptomssymptomsheart-arrhythmias-symptomssymptoms
