datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatDoctor-HealthCareMagic-100k
Dataset Card for "ChatDoctor-HealthCareMagic-100k"
More Information needed
ChatDoctor-iCliniq
Dataset Card for "ChatDoctor-iCliniq"
More Information needed
ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1ChatDoctor_HealthCareMagicThe ChatDoctor-HealthCareMagic-100k dataset comprises 112,000 real-world medical question-and-answer pairs, providing a substantial and diverse collection of authentic medical dialogues. There is a slight risk to this dataset since there are grammatical inconsistencies in many of the questions and answers, but this can potentially help separate strong healthcare retrieval models from weak ones.
Usage
import datasets
# Download the dataset
queries =… See the full description on the dataset page: https://huggingface.co/datasets/embedding-benchmark/ChatDoctor_HealthCareMagic.chatdoctor-200kThis ChatDoctor-200K dataset is collected from this paper https://arxiv.org/pdf/2303.14070.pdf
Alternatively, you can download the original dataset from this link https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view?usp=sharing
lavita-ChatDoctor-HealthCareMagic-100kchat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.chatdoctor-embedded
Chat Doctor with Embeddings
This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped:
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
414k
Token Count
1.7b
Origin
https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view
Source of raw data
?
Processing details
paper
Embedding Model
BAAI/bge-small-en-v1.5
Data Diversity
index
Example Output
GPT-4 Rationale
GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.chatdoctor-5kThis ChatDoctor-5K dataset is collected from this paper https://arxiv.org/pdf/2303.14070.pdf
Alternatively, you can download the original dataset from this link https://drive.google.com/file/d/1nDTKZ3wZbZWTkFMBkxlamrzbNz0frugg/view?usp=sharing
chatdoctor200k
Dataset Card for "chatdoctor200k"
More Information needed
chatdoctor-healthcaremagic-112k-vi
🩺 ChatDoctor HealthCareMagic 112k (Vietnamese Translated)
Tập dữ liệu hỏi đáp y khoa ChatDoctor HealthCareMagic 112k được dịch sang tiếng Việt chất lượng cao, phục vụ fine-tune các mô hình ngôn ngữ lớn (LLM) trong lĩnh vực y tế, chăm sóc sức khỏe và tư vấn y khoa tổng quát.
📌 Tổng quan dữ liệu
Quy mô: 112,165 cặp hỏi - đáp y tế thực tế giữa bệnh nhân và bác sĩ.
Phân chia:
train: 106,556 mẫu (95%)
validation: 5,609 mẫu (5%)
Định dạng: Chuẩn Alpaca /… See the full description on the dataset page: https://huggingface.co/datasets/NoirHuy/chatdoctor-healthcaremagic-112k-vi.ChatDoctor-RL
Intelligent-Internet/ChatDoctor-Improved-Answer Dataset
This dataset represents a carefully curated subset derived from the original ChatDoctor-HealthCareMagic-100k[lavita/ChatDoctor-HealthCareMagic-100k] dataset, where we have undertaken significant improvements to enhance the quality and depth of the responses. The answers have been thoroughly refined to provide greater detail, clarity, and precision, while incorporating a heightened focus on safety awareness to ensure responsible… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/ChatDoctor-RL.rteb-ChatDoctorRetrieval
ChatDoctorRetrieval — RTEB open subset, unified schema
A normalised copy of the dataset behind the mteb task ChatDoctorRetrieval, one of the 17 open tasks in the
RTEB(beta) retrieval benchmark. Same queries, documents and relevance
judgements as the benchmark evaluates — reshaped into one strict schema shared by all 17.
Source
embedding-benchmark/ChatDoctor_HealthCareMagic @ 50c2986fedff (the revision pinned in mteb)
Domain · languages
healthcare · eng
Queries /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/rteb-ChatDoctorRetrieval.ChatDoctor-iCliniq-7.3k
Dataset Card for "ChatDoctor-iCliniq-7.3k"
More Information needed
chatdoctor-200k-stripped-embeddedChatDoctor_HealthCareMagic_112k
Dataset Card for "ChatDoctor_HealthCareMagic_112k"
More Information needed
chatdoctorChatDoctorRetrieval
ChatDoctorRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
A medical retrieval task based on ChatDoctor_HealthCareMagic dataset containing 112,000 real-world medical question-and-answer pairs. Each query is a medical question from patients (e.g., 'What are the symptoms of diabetes?'), and the corpus contains medical responses and healthcare information. The task is to retrieve the correct medical information that answers the patient's question. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ChatDoctorRetrieval.ChatDoctor-HealthCareMagic-100k
Dataset Card for "ChatDoctor-HealthCareMagic-100k"
More Information needed
ChatDoctor_chatdoctor_7k
Dataset Card for "ChatDoctor_chatdoctor_7k"
数据集名称:lavita/ChatDoctor-iCliniq
数据集原型来源:https://huggingface.co/datasets/lavita/ChatDoctor-iCliniq
数据规模:7.32k
数据生成:由llm生成
数据领域:医患对话
More Information needed
chatdoctor_refChatDoctor_chatGpt_7k
Dataset Card for "ChatDoctor_chatGpt_7k"
More Information needed
cleaned_chatdoctor_healthcaremagic_100kMedical-ChatDoctor-HealthCareMagic-100k-Qwen3-verifyChatDoctor-50k-qaMedical-ChatDoctor-HealthCareMagic-100k-Qwen3ChatDoctor-HealthCareMagic-100k-Diabetes
Dataset Card for "ChatDoctor-HealthCareMagic-100k"
More Information needed
chatdoctor-datasetChatDoctor-HealthCareMagic-100k-fixedChatDoctor-HealthCareMagic-100k-clean
