CoolFace
Datasetpublic

faur-ai/ro-chatdoctor-200k

This dataset is the translated ChatDoctor-200k instruct dataset using LLMic, a bilingual Romanian-English LLM. ChatDoctor-200k is a large dataset of 100,000 patient-doctor dialogues sourced from a widely used online medical consultation platform. @article{li2023chatdoctor, title={Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge}, author={Li, Yunxiang and Li, Zihan and Zhang, Kai and Dan, Ruilong and Jiang, Steve and Zhang… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-chatdoctor-200k.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes62downloads
Dataset Card

This dataset is the translated ChatDoctor-200k instruct dataset using LLMic, a bilingual Romanian-English LLM.

ChatDoctor-200k is a large dataset of 100,000 patient-doctor dialogues sourced from a widely used online medical consultation platform.

@article{li2023chatdoctor, title={Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge}, author={Li, Yunxiang and Li, Zihan and Zhang, Kai and Dan, Ruilong and Jiang, Steve and Zhang, You}, journal={Cureus}, volume={15}, number={6}, year={2023}, publisher={Cureus} }

@article{buadoiu2025llmic,
  title={LLMic: Romanian Foundation Language Model},
  author={B{\u{a}}doiu, Vlad-Andrei and Dumitru, Mihai-Valentin and Gherghescu, Alexandru M and Agache, Alexandru and Raiciu, Costin},
  journal={arXiv preprint arXiv:2501.07721},
  year={2025}
}