datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doctor-patient-conversations-3000💼 Commercial Use License
This dataset is free for research use (CC-BY-NC-4.0).
For commercial use, AI model training inside products, or enterprise usage:
👉 License fee: $49
📩 Contact: syntech.ai.official@gmail.com
license: cc-by-nc-4.0
task_categories:
- text-classification
language:
- en
tags:
- medical
- synthetic-data
- healthcare
- doctor-patient
- conversations
- ai-dataset
- llm-training
- jsonl
- csv
📘 Doctor–Patient Synthetic Conversation Dataset (3,000 Samples)
A… See the full description on the dataset page: https://huggingface.co/datasets/syntech-ai/doctor-patient-conversations-3000.chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.no-robots-sharegpt
no-robots-sharegpt
HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information.
no-robots-sharegpt.jsonl
Original dataset converted to ShareGPT
no-robots-sharegpt-fixed.jsonl
Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.Doctor-patient-convo
AI-Generated Doctor-Patient Conversations & Clinical Summaries
Dataset Description
This dataset contains synthetic, AI-generated transcripts of doctor-patient conversations paired with summarized clinical notes detailing the medical outcomes. It is designed to train and fine-tune Small Language Models (SLMs) for medical text summarization tasks.
Data Generation Methodology
The data was entirely generated using Google's Gemma 3 27B model.
To ensure… See the full description on the dataset page: https://huggingface.co/datasets/Tser-vak/Doctor-patient-convo.capybara-sharegpt
capybara-sharegpt
LDJnr/Capybara converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information. All credit goes to the original creator.
theory-of-mind-dpoThis is grimulkan/theory-of-mind with "rejected" responses generated using mistralai/Mistral-7B-Instruct-v0.2, and the file formatted for use in DPO training.
The code used to generate the dataset can be found in this repository: https://github.com/DocShotgun/LLM-datagen
bonsai_8b_distilled_edited_106byDoctorEdoP369🧠 [bonsai_8b_distilled_edited_106byDoctorEdoP369 ]
A highly curated, gold-standard dataset of 106 refined reasoning traces built for fine-tuning compact AI models.
This dataset contains 106 high-density examples specifically designed for complex Chain of Thought (CoT) reasoning.
Every single entry was thoroughly cleaned, mathematically verified, and enhanced through an advanced refining pipeline guided by Claude Opus 4.8 and Gemini Pro Extended.
Zero Logical & Mathematical Errors: All… See the full description on the dataset page: https://huggingface.co/datasets/DoctorEdoP369/bonsai_8b_distilled_edited_106byDoctorEdoP369.n-gramsDoctor-Shotgun_theory-of-mind-dpo-PreferenceShareGPTkalo-opus-misc-kto-combineddoctor-patient-conversation-v1kalo-22k-norefusal-kto-combineddoctorai-datasetChat_Doctorkalo-3k-filtered-kto-combinedchat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/JATINGAHLOT/chat_doctor.synthstruct-1.1-kto-combinedchat_doctor-qa-pashto
Chat Doctor QA Pashto
Short medical Q&A dataset translated from Indonesian → Pashto.Useful for Pashto medical assistants, clinical reasoning SFT, and healthcare chatbots.
Dataset Summary
This dataset contains short doctor–patient style questions and answers.Each entry follows the Ministral‑Instruct chat format:
{
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
]
}
Languages
ps — Pashto
Source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/chat_doctor-qa-pashto.BestMathQualityByDoctorEdoP369.jsonl
A highly curated, gold-standard dataset for fine-tuning AI models on complex mathematical reasoning.
Finding truly high-quality math data is a long and difficult process. Most open-source math datasets are noisy, poorly formatted, or worse—packed with actual calculation errors that degrade your model's performance.
This dataset solves problems. I took the absolute best math datasets available, unified them into a single file (BestMathQualityByDoctorEdoP369.jsonl). Quality by… See the full description on the dataset page: https://huggingface.co/datasets/DoctorEdoP369/BestMathQualityByDoctorEdoP369.jsonl.DoctordoctorPersonacdndoctorchat-doctor-qa-pashtoDOctor_Aidoctor-answerdoctor-dataset-updateddoctor-dataset-districthat_doctor2_tr_200kdoctor1_2412
