CoolFace
Datasetpublic

Lyncia/LingxiDiag-16K

LingxiDiag-16K A Large-Scale Synthetic Psychiatric Dialogue Dataset for Diagnostic Decision Support Overview LingxiDiag-16K is a synthetic psychiatric dialogue dataset containing approximately 16,000 electronic medical records (EMRs) and doctor-patient consultation dialogues. The dataset is designed for evaluating and training LLM-based psychiatric diagnostic decision support systems, with demographically aligned distributions reflecting real-world clinical… See the full description on the dataset page: https://huggingface.co/datasets/Lyncia/LingxiDiag-16K.

sourceHugging Facecc-by-nc-4.0updated 8mo agoView on Hugging Face
2likes95downloads
Dataset Card

LingxiDiag-16K

A Large-Scale Synthetic Psychiatric Dialogue Dataset for Diagnostic Decision Support

Overview

LingxiDiag-16K is a synthetic psychiatric dialogue dataset containing approximately 16,000 electronic medical records (EMRs) and doctor-patient consultation dialogues. The dataset is designed for evaluating and training LLM-based psychiatric diagnostic decision support systems, with demographically aligned distributions reflecting real-world clinical settings.

This dataset is part of the LingxiDiagBench project.

Dataset Details

  • —Size: ~16,000 samples
  • —Language: Chinese
  • —Domain: Psychiatry / Mental Health
  • —License: CC BY-NC 4.0

Data Fields

FieldDescription
patient_idUnique patient identifier
AgePatient age
GenderPatient gender
DepartmentClinical department
ChiefComplaintChief complaint
PresentIllnessHistoryHistory of present illness
PersonalHistoryPersonal history
FamilyHistoryFamily history
DiagnosisClinical diagnosis
DiagnosisCodeICD-10 diagnosis code
cleaned_textDoctor-patient dialogue text
icd_clf_labelICD classification label
two_class_labelBinary classification label
four_class_label4-class classification label

Splits

SplitSamples
Train~14000
Validation~1000
Test~1000

Diagnostic Categories

  • —2-class: Depression, Anxiety
  • —4-class: Depression, Anxiety, Mixed, Others
  • —12-class: F20, F31, F32, F39, F41, F42, F43, F45, F51, F98, Z71, Others

Usage

python
from datasets import load_dataset

dataset = load_dataset("XuShihao6715/LingxiDiag-16K")

Citation

If you use this dataset in your research, please cite:

bibtex
@article{lingxidiagbench2026,
  title={LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis},
  author={Shihao Xu et al.},
  journal={arXiv preprint},
  year={2026}
}

License

This dataset is licensed under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).

You are free to:

  • —Share — copy and redistribute the material in any medium or format
  • —Adapt — remix, transform, and build upon the material

Under the following terms:

  • —Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made.
  • —NonCommercial — You may not use the material for commercial purposes.

Contact

Made by the Evermind Lingxi Team from Shanda Group.

Join us: https://evermind.ai/careers