datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.TCM-Ladder
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
TCM-Ladder is the first multimodal QA dataset specifically designed for evaluating large TCM language models. The dataset spans multiple core disciplines of TCM, including fundamental theory, diagnostics, herbal formulas, internal medicine, surgery, pharmacognosy, and pediatrics.
Dataset Details
Dataset Description
TCM-Ladder is the first multimodal QA dataset specifically… See the full description on the dataset page: https://huggingface.co/datasets/timzzyus/TCM-Ladder.TCMQA
TCMQA
Built by TechTCM.
A Traditional Chinese Medicine question-answering benchmark: 33,872 Chinese-language
multiple-choice questions drawn from TCM licensing-exam material, plus a
5,250-question subset answered by 102 licensed TCM practitioners, giving a human
reference point for the same items a model is scored on.
Configs
qa (default) — 33,872 rows
Field
Type
Notes
id
int64
Stable question id. Not contiguous — ids are preserved from… See the full description on the dataset page: https://huggingface.co/datasets/TechTCM/TCMQA.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.tcm-divination-training
TCM & Divination Training Dataset v2
Comprehensive training dataset for Bazi, Tử Vi (Zi Wei Dou Shu), TCM, and divination domains.
Dataset Summary
Metric
Value
Total Unique Samples
162,384
File Size
651 MB
Languages
Vietnamese, English, Chinese
Last Updated
2026-01-11
Data Sources
Source
Unique Samples
Description
bazi_books
74,533
Extracted from Bazi/Tử Vi books (OCR)
gpt_training_ready
48,551
GPT-generated Q&A pairs… See the full description on the dataset page: https://huggingface.co/datasets/jakeveo05/tcm-divination-training.TCM-Ladder
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
TCM-Ladder is the first multimodal QA dataset specifically designed for evaluating large TCM language models. The dataset spans multiple core disciplines of TCM, including fundamental theory, diagnostics, herbal formulas, internal medicine, surgery, pharmacognosy, and pediatrics.
Dataset Details
Dataset Description
TCM-Ladder is the first multimodal QA dataset specifically… See the full description on the dataset page: https://huggingface.co/datasets/Xuqiang2026/TCM-Ladder.
