datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.ShenNong_TCM_Dataset2023_Pharmacist_Licensure_Examination-TCM_trackThe 2023 Chinese National Pharmacist Licensure Examination is divided into two distinct tracks: the Pharmacy track and the Traditional Chinese Medicine (TCM) Pharmacy track. The data provided here pertains to the Traditional Chinese Medicine (TCM) Pharmacy track examination. It is important to note that this dataset was collected from online sources, and there may be some discrepancies between this data and the actual examination.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/2023_Pharmacist_Licensure_Examination-TCM_track.TCMLE
中医执业医师资格考试题库数据集 📚
数据来源
国家执业医师资格考试 中医执业医师考试 真题
国家执业医师资格考试 中医执业医师考试 模拟题
国家执业医师资格考试 中医执业助理医师考试 真题
国家执业医师资格考试 中医执业助理医师考试 模拟题
数据规模 📊
本题库共有7956道题目。
其中:
👩⚕️ 助理医师题目: 2700道
👨⚕️ 执业医师题目: 5256道
题型结构与数量 🏗️
📁 /
│
├── 📁 Assistant/ # 中医执业助理医师
│ ├── 📁 Theory_Questions/ # 理论题目
│ │ ├── 📁 Year_1/
│ │ │ ├── Mock.json # 222题
│ │ │ └── Past_Paper.json # 245题
│ │ ├── 📁 Year_2/… See the full description on the dataset page: https://huggingface.co/datasets/Bolin97/TCMLE.EnTubeTCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.TCM-Ancient-Modern-Open
TCM Ancient-to-Modern Chinese Open Metadata
Chinese documentation | Code repository
This public metadata companion was designed to avoid redistribution of third-party published text from the controlled TCM Ancient-to-Modern Chinese parallel corpus. It releases the complete 9,610-record index, frozen split assignments, source-edition register, length metadata, cryptographic integrity digests, and ten readable demonstration pairs. It does not redistribute the experimental… See the full description on the dataset page: https://huggingface.co/datasets/xs12345/TCM-Ancient-Modern-Open.tcm-divination-training
TCM & Divination Training Dataset v2
Comprehensive training dataset for Bazi, Tử Vi (Zi Wei Dou Shu), TCM, and divination domains.
Dataset Summary
Metric
Value
Total Unique Samples
162,384
File Size
651 MB
Languages
Vietnamese, English, Chinese
Last Updated
2026-01-11
Data Sources
Source
Unique Samples
Description
bazi_books
74,533
Extracted from Bazi/Tử Vi books (OCR)
gpt_training_ready
48,551
GPT-generated Q&A pairs… See the full description on the dataset page: https://huggingface.co/datasets/jakeveo05/tcm-divination-training.tcm-collected-works
Collected Works & Treatises · 中医医集医论 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Masters' complete works + medical treatises/discourse: 名家集著, 医理集论, 医话
171 public-domain works of classical Traditional Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-collected-works.tcm-case-records
Case Records · 中医医案 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Classical clinical case records: 名医类案·临证指南医案·古今医案 (医案)
33 public-domain works of classical Traditional Chinese Medicine, as clean full text… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-case-records.TCM_SFT_datasettcm-reference-compendia
Reference Compendia · 中医类书全录 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Encyclopedic compendia (few works, huge volume): 古今图书集成医部全录(123卷)·医方类聚·四库医家类
14 public-domain works of classical Traditional Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-reference-compendia.tcm-formulary
Formulary · 中医方书 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Classical prescription collections: 局方·千金方·外台秘要·医方集解 (方剂)
91 public-domain works of classical Traditional Chinese Medicine, as clean full text… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-formulary.tcm-diagnostics
Diagnostics · 中医诊法脉学 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Pulse/tongue/inspection + pattern diagnosis: 脉经·濒湖脉学·舌鉴 (脉学·辩证诊治)
42 public-domain works of classical Traditional Chinese Medicine, as clean full… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-diagnostics.tc_medical_base
tc_medical_base — 中醫典籍檢索語料
Clause-level (條文) corpus of five classical TCM texts with precomputed
sentence-BERT embeddings, built for the 典籍提示 (classical-text hint)
retrieval feature of the tcm-homevisit 中醫居家醫療 app (Taiwan).
Books (retrieval units after parsing)
書名
clauses
傷寒論(宋本)
603
傷寒雜病論(桂林古本)
995
金匱玉函經
1,215
金匱要略方論
590
傅青主女科
166
total
3,569
(古本康平傷寒論 was included in an earlier revision and removed 2026-08: its
clauses overlap heavily… See the full description on the dataset page: https://huggingface.co/datasets/huckiyang/tc_medical_base.tcm-gynecology-pediatrics
Gynecology & Pediatrics · 中医妇产儿科 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Women's & children's medicine: 妇科·产科·儿科·广嗣 (傅青主女科, 幼科 etc.)
69 public-domain works of classical Traditional Chinese Medicine, as clean… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-gynecology-pediatrics.tcm-acupuncture-classics
Acupuncture & Channels · 中医针灸经络 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Pre-modern acupuncture/moxa + channels: 针灸甲乙经·针灸大成·铜人腧穴 (针灸·经络)
33 public-domain works of classical Traditional Chinese Medicine, as… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-acupuncture-classics.TCM-m3-SFT-datasettcm-materia-medica
Materia Medica · 中医本草 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Herbal drug monographs + processing/properties: 神农本草经·本草纲目·本草崇原 (本草·炮炙·药性)
59 public-domain works of classical Traditional Chinese Medicine, as… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-materia-medica.tcm-external-surgical
External & Surgical Medicine · 中医外科疡科 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Surgery, sores, ENT/limbs, deficiency: 外科·疡疹痧痘·五官四肢·风劳虚损·外治推拿
50 public-domain works of classical Traditional Chinese Medicine… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-external-surgical.tcm-health-cultivation
Health Cultivation · 中医养生 💰 (Commercial Dataset)
This is a commercial dataset. A free 3-work sample is provided below; the
full dataset is available for licensing/purchase.
📧 To purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — derived from public-domain classical works.
What you get
Classical 养生 / longevity texts: 养性延命录(陶弘景)·养生导引法·养生类要
18 public-domain works of classical Traditional Chinese Medicine, as clean full text… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/tcm-health-cultivation.ChatMed_TCM-gemma4-10000Baize-TCM-Corpus-for-Large-Language-Models-V2
白泽中医药大模型语料库
版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.Baize-TCM-Corpus-for-Large-Language-Models-V1
白泽中医药大模型语料库
版本:1.0语料数量:4,735 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 4,735 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V1.ShenNong_TCM_DatasetTCM_19W_DataSet-SFT数据介绍
非网络来源的高质量指令微调中医数据集(部分内容)
任何问题请联系:longfeichai@stu.haust.edu.cn
TCMD-EvalTCM_19W_DataSet-SFT数据介绍
非网络来源的高质量指令微调中医数据集(部分内容)
任何问题请联系:longfeichai@stu.haust.edu.cn
