datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HealthChat-11K
HealthChat-11K
This repository contains HealthChat-11K, a curated dataset of approximately 11,000 real-world conversations, composed of 25,000 user messages, where users seek healthcare information from Large Language Models (LLMs). The goal of this work is to provide a high-quality resource for systematically studying and improving health conversations involving humans and AI (e.g., LLMs).
The dataset was presented in the paper: "What's Up, Doc?": Analyzing How Users Seek Health… See the full description on the dataset page: https://huggingface.co/datasets/yahskapar/HealthChat-11K.health-chatbot
Dataset Card for Dataset Name
Health Question and Answer Clean Dataset
Dataset Details
Dataset Description
This dataset provides a detailed overview of health question & answer pairs. It includes data on health problems and corresponding answers, making it suitable for variable tasks like healthcare chatbot training.
Language(s) (NLP): English
License: Apache-2.0
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/shaneperry0101/health-chatbot.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.bangla-health-related-paraphrased-dataset
Dataset Card for "BanglaHealthParaphrase"
BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.mental_health_counseling_conversations_rated
Dataset Card for Mental Health Counseling Conversations Rated
This dataset extends the existing dataset Mental Health Counseling Conversations and adds ratings for the responses.
Dataset Details
This dataset is an extension for the dataset Mental Health Counseling Conversations.
It adds ratings for the responses generated by four different LLMs. The responses are rated across the following dimensions:
empathy
appropriateness
relevance
The following four LLMs are used… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_conversations_rated.Global_Health-Nutrition-And-Population-Statistics
Global_Health-Nutrition-And-Population-Statistics Dataset
This Dataset contains all verified and authorized Health, Nutrition and Population Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Health-Nutrition-And-Population-Statistics.mental_health_counseling_responses
Dataset Card for Mental Health Counseling Responses
This dataset contains responses to questions from mental health counseling sessions.
The responses are rated by LLMs using the dimensions: empathy, appropriateness, and relevance.
A detailed explanation of the rating process can be found in this blog post.
For a detailed analysis of LLM-generated responses and their comparison to human responses, refer to this blog post.
The original data with the human responses can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_responses.spai-ss6-corpus-mental-health-thai
SPAI SS6 Thai Mental Health Index
Index repo for the imported Thai mental-health dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: mental_health_thai
Rows in canonical config: 21,544
Parquet size in canonical config: 0.02 GB
Source license: unknown
License review… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-mental-health-thai.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.tw-health-43M
Dataset Card for tw-health-43M
tw-health-43M 是一個台灣健康醫療領域的繁體中文預訓練語料集,包含 48,627 篇文章,總計約 4,366 萬 tokens。涵蓋疾病衛教、醫療新聞、公共衛生政策與健康促進等內容,適用於語言模型在醫療健康領域之持續預訓練。
Dataset Details
Dataset Description
本資料集彙整台灣健康醫療相關之文章,涵蓋內科、外科、兒科、精神科、營養、復健、傳染病防治等多元醫療主題。每篇文章附帶 token 數、字數、來源 URL 與更新日期等元資料。
Curated by: Liang Hsun Huang
Language(s) (NLP): Traditional Chinese
License: CC BY-NC-SA 4.0
Dataset Sources
Repository: lianghsun/tw-health-43M… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-health-43M.Team_Health_Diagnosing_Dysfunction_Theory
Team Health Diagnosing Dysfunction — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Team_Health_Diagnosing_Dysfunction_Theory.spai-ss6-corpus-medical-health-web
SPAI SS6 Thai Medical Health Web Corpus
Thai public medical and health web articles collected by the local scraping pipeline.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: default
Rows in canonical config: 3,660
Parquet size in canonical config: 0.01 GB
Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.dummy_health_data
Synthetic Healthcare Dataset
Overview
This dataset is a synthetic healthcare dataset created for use in data analysis. It mimics real-world patient healthcare data and is intended for applications within the healthcare industry.
Data Generation
The data has been generated using the Faker Python library, which produces randomized and synthetic records that resemble real-world data patterns. It includes various healthcare-related fields such as patient… See the full description on the dataset page: https://huggingface.co/datasets/vrajakishore/dummy_health_data.Team_Health_Diagnosing_Dysfunction_Practical
Team Health Diagnosing Dysfunction — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Team_Health_Diagnosing_Dysfunction_Practical.
