CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syntech-ai /doctor-patient-conversations-3000💼 Commercial Use License This dataset is free for research use (CC-BY-NC-4.0). For commercial use, AI model training inside products, or enterprise usage: 👉 License fee: $49 📩 Contact: syntech.ai.official@gmail.com license: cc-by-nc-4.0 task_categories: - text-classification language: - en tags: - medical - synthetic-data - healthcare - doctor-patient - conversations - ai-dataset - llm-training - jsonl - csv 📘 Doctor–Patient Synthetic Conversation Dataset (3,000 Samples) A… See the full description on the dataset page: https://huggingface.co/datasets/syntech-ai/doctor-patient-conversations-3000.text1K<n<10K1 likes283 downloads10mo agoHugging Face02avaliev /chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work. 100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED 10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED 5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually) data sample: {'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.textquestion-answering100K<n<1M15 likes236 downloads3y agoHugging Face03Doctor-Shotgun /no-robots-sharegpt no-robots-sharegpt HuggingFaceH4/no_robots with both test and train splits combined and converted to ShareGPT format for use in common training repositories. Please refer to the original repository's dataset card for more information. no-robots-sharegpt.jsonl Original dataset converted to ShareGPT no-robots-sharegpt-fixed.jsonl Manual edits were made to ~10 dataset entries that were throwing warnings in axolotl - turns out that some of the multi-turn conversations had… See the full description on the dataset page: https://huggingface.co/datasets/Doctor-Shotgun/no-robots-sharegpt.texttext-generation10K<n<100K24 likes133 downloads3y agoHugging Face04Tser-vak /Doctor-patient-convo AI-Generated Doctor-Patient Conversations & Clinical Summaries Dataset Description This dataset contains synthetic, AI-generated transcripts of doctor-patient conversations paired with summarized clinical notes detailing the medical outcomes. It is designed to train and fine-tune Small Language Models (SLMs) for medical text summarization tasks. Data Generation Methodology The data was entirely generated using Google's Gemma 3 27B model. To ensure… See the full description on the dataset page: https://huggingface.co/datasets/Tser-vak/Doctor-patient-convo.textn<1K1 likes111 downloads13d agoHugging Face05Doctor-Shotgun /capybara-sharegpt capybara-sharegpt LDJnr/Capybara converted to ShareGPT format for use in common training repositories. Please refer to the original repository's dataset card for more information. All credit goes to the original creator. texttext-generation10K<n<100K4 likes51 downloads3y agoHugging Face06Doctor-Shotgun /theory-of-mind-dpoThis is grimulkan/theory-of-mind with "rejected" responses generated using mistralai/Mistral-7B-Instruct-v0.2, and the file formatted for use in DPO training. The code used to generate the dataset can be found in this repository: https://github.com/DocShotgun/LLM-datagen textn<1K18 likes41 downloads3y agoHugging Face07DoctorEdoP369 /bonsai_8b_distilled_edited_106byDoctorEdoP369🧠 [bonsai_8b_distilled_edited_106byDoctorEdoP369 ] A highly curated, gold-standard dataset of 106 refined reasoning traces built for fine-tuning compact AI models. This dataset contains 106 high-density examples specifically designed for complex Chain of Thought (CoT) reasoning. Every single entry was thoroughly cleaned, mathematically verified, and enhanced through an advanced refining pipeline guided by Claude Opus 4.8 and Gemini Pro Extended. Zero Logical & Mathematical Errors: All… See the full description on the dataset page: https://huggingface.co/datasets/DoctorEdoP369/bonsai_8b_distilled_edited_106byDoctorEdoP369.textn<1K0 likes33 downloads2mo agoHugging Face08Doctor-Bob /n-gramstext1M<n<10M0 likes27 downloads10d agoHugging Face09PJMixers /Doctor-Shotgun_theory-of-mind-dpo-PreferenceShareGPTtextreinforcement-learningn<1K1 likes20 downloads2y agoHugging Face10Doctor-Shotgun /kalo-opus-misc-kto-combinedtext1K<n<10K1 likes19 downloads2y agoHugging Face11okyashgajjar /doctor-patient-conversation-v1text1K<n<10K1 likes18 downloads5mo agoHugging Face12Doctor-Shotgun /kalo-22k-norefusal-kto-combinedtext10K<n<100K2 likes13 downloads2y agoHugging Face13dimitars /doctorai-datasettextn<1K1 likes12 downloads3y agoHugging Face14xsshacker /Chat_Doctortext1K<n<10K0 likes12 downloads2y agoHugging Face15Doctor-Shotgun /kalo-3k-filtered-kto-combinedtext1K<n<10K1 likes12 downloads2y agoHugging Face16JATINGAHLOT /chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work. 100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED 10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED 5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually) data sample: {'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/JATINGAHLOT/chat_doctor.textquestion-answering100K<n<1M0 likes12 downloads6mo agoHugging Face17Doctor-Shotgun /synthstruct-1.1-kto-combinedtext10K<n<100K2 likes11 downloads2y agoHugging Face18nassimjp /chat_doctor-qa-pashto Chat Doctor QA Pashto Short medical Q&A dataset translated from Indonesian → Pashto.Useful for Pashto medical assistants, clinical reasoning SFT, and healthcare chatbots. Dataset Summary This dataset contains short doctor–patient style questions and answers.Each entry follows the Ministral‑Instruct chat format: { "messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ] } Languages ps — Pashto Source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/chat_doctor-qa-pashto.text1K<n<10K0 likes11 downloads3mo agoHugging Face19DoctorEdoP369 /BestMathQualityByDoctorEdoP369.jsonlgated A highly curated, gold-standard dataset for fine-tuning AI models on complex mathematical reasoning. Finding truly high-quality math data is a long and difficult process. Most open-source math datasets are noisy, poorly formatted, or worse—packed with actual calculation errors that degrade your model's performance. This dataset solves problems. I took the absolute best math datasets available, unified them into a single file (BestMathQualityByDoctorEdoP369.jsonl). Quality by… See the full description on the dataset page: https://huggingface.co/datasets/DoctorEdoP369/BestMathQualityByDoctorEdoP369.jsonl.text1K<n<10K0 likes9 downloads1mo agoHugging Face20Henil1 /Doctortext100K<n<1M2 likes8 downloads3y agoHugging Face21tinycrops /doctorPersonatabular10K<n<100K0 likes8 downloads2y agoHugging Face22DoctorSlimm /cdnimagen<1K0 likes8 downloads5mo agoHugging Face23dipesh1111 /doctortextn<1K0 likes6 downloads4y agoHugging Face24nassimjp /chat-doctor-qa-pashtotextn<1K0 likes6 downloads3mo agoHugging Face25avijitbhuin21 /DOctor_Aitext100K<n<1M0 likes5 downloads2y agoHugging Face26libin46 /doctor-answertext1K<n<10K1 likes4 downloads3y agoHugging Face27IgYahiko /doctor-dataset-updatedtext1M<n<10M0 likes4 downloads2y agoHugging Face28IgYahiko /doctor-dataset-districttext1M<n<10M0 likes4 downloads2y agoHugging Face29falan42 /hat_doctor2_tr_200ktext100K<n<1M0 likes3 downloads2y agoHugging Face30ssunny /doctor1_2412textn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.