datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Diabetes-Clinical-Intruction-ENGGTHUB LINK: (https://github.com/Bernardosalerno/Diabetes-Clinical-Instruction-Dataset-LLM-SFT-FineTuning-JSONL)
💎 Dataset Overview
Diabetes-Clinical-Instruction-ENG is a premium, medically-validated instruction-tuning dataset containing 745 high-quality samples. This page contains a preview of 20 samples. The full dataset is focused on diabetology, glycemic management, and patient care.
🚀 FOR ALL THE SAMPLES: The first 50 copies are available for just €25. After that, the price… See the full description on the dataset page: https://huggingface.co/datasets/Bernardosalerno/Diabetes-Clinical-Intruction-ENG.ada_diabetes_5000_instruction
ADA Diabetes Instruction Dataset (5,000 Samples)
This dataset contains 5,000 synthetic yet clinically-informed patient cases for Type 2 diabetes, designed for instruction tuning of language models (e.g., Gemma 3, Unsloth) to recommend ADA guideline-based therapies with drug-specific dosing.
Dataset Overview
Task: Given a patient profile, recommend ADA-aligned diabetes treatment including therapy, drug-specific starting doses, and rationale.
Size: 5,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/ada_diabetes_5000_instruction.ehr-diabetes-cohort-100
HipAAsynth Dataset
Summary
This dataset is a validation artifact generated by HipAAsynth.
HipAAsynth is a deterministic testing and validation service that simulates real-world variability to evaluate how healthcare systems perform under deployment conditions.
Description
This dataset represents a controlled cohort used for testing and benchmarking.
HipAAsynth generates cohorts to simulate how conditions present across:
patient populations
demographic… See the full description on the dataset page: https://huggingface.co/datasets/HipAAsynth/ehr-diabetes-cohort-100.
