datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HealthChat-11K
HealthChat-11K
This repository contains HealthChat-11K, a curated dataset of approximately 11,000 real-world conversations, composed of 25,000 user messages, where users seek healthcare information from Large Language Models (LLMs). The goal of this work is to provide a high-quality resource for systematically studying and improving health conversations involving humans and AI (e.g., LLMs).
The dataset was presented in the paper: "What's Up, Doc?": Analyzing How Users Seek Health… See the full description on the dataset page: https://huggingface.co/datasets/yahskapar/HealthChat-11K.healthbench-psych
HealthBench-Psych
An expert-adjudicated mental-health subset of HealthBench
(OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations), together with
released results for 23 language models under a three-judge panel.
Maintained by MindBench.ai · Division of Digital Psychiatry, Beth Israel
Deaconess Medical Center. Code and evaluation harness:
github.com/mindbench-ai/healthbench-psych.
Subset
n
Definition
healthbench-psych-v2
611… See the full description on the dataset page: https://huggingface.co/datasets/mindbench-ai/healthbench-psych.Health_Information_Seeking_under_Limited_Evidence
Health Information Seeking under Limited Evidence (HISLE)
HISLE is a clinically informed benchmark for evaluating LLM-based agents responding to incomplete mental-health information needs.
File
Records
Contents
matched_pairs_47.jsonl
47
Matched Chinese–English scenario pairs
matched_variants_3290.jsonl
3,290
Query variants for the matched scenarios
coverage_originals_24.jsonl
24
Coverage-expansion queries
coverage_variants_840.jsonl
840
Query variants for… See the full description on the dataset page: https://huggingface.co/datasets/PsychiatryAgentBench25/Health_Information_Seeking_under_Limited_Evidence.HealthyLife-Insurance-Charge-Prediction-v2tamil-nadu-government-health-facilities
Tamil Nadu Government Health Facilities
A cleaned and structured dataset of government healthcare facilities across Tamil Nadu, India.
Dataset Description
This dataset contains 2,398 healthcare facility records across Tamil Nadu, including government hospitals, Primary Health Centres (PHCs), Urban Primary Health Centres (UPHCs), maternity facilities, dental facilities, and other government healthcare facilities.
Each record includes information such as:
Facility… See the full description on the dataset page: https://huggingface.co/datasets/lildosa/tamil-nadu-government-health-facilities.m42-health__Llama3-Med42-70B-details
Dataset Card for Evaluation run of m42-health/Llama3-Med42-70B
Dataset automatically created during the evaluation run of model m42-health/Llama3-Med42-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/m42-health__Llama3-Med42-70B-details.tamil-nadu-government-health-facilities
Tamil Nadu Government Health Facilities
A cleaned and structured dataset of government healthcare facilities across Tamil Nadu, India.
Dataset Description
This dataset contains 2,398 healthcare facility records across Tamil Nadu, including government hospitals, Primary Health Centres (PHCs), Urban Primary Health Centres (UPHCs), maternity facilities, dental facilities, and other government healthcare facilities.
Each record includes information such as:
Facility… See the full description on the dataset page: https://huggingface.co/datasets/naveenpixel/tamil-nadu-government-health-facilities.han-humanoid-battery-health-logs-v1
Humanoid Battery Health Logs (HBHL)
Description
Battery performance data collected from humanoid
robot operational cycles.
Features
cycle_id (string)
charge_cycles_count (int)
average_temperature_c (float)
discharge_rate (float)
voltage_variation (float)
battery_health_status (low/medium/high)
Target
battery_health_status
Use Cases
Battery degradation modeling
Maintenance scheduling
Risk prediction
Evaluation Metrics… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-battery-health-logs-v1.mental_health_counseling_conversations_rated
Dataset Card for Mental Health Counseling Conversations Rated
This dataset extends the existing dataset Mental Health Counseling Conversations and adds ratings for the responses.
Dataset Details
This dataset is an extension for the dataset Mental Health Counseling Conversations.
It adds ratings for the responses generated by four different LLMs. The responses are rated across the following dimensions:
empathy
appropriateness
relevance
The following four LLMs are used… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_conversations_rated.health-log-extraction-datasetTo create this dataset, we sampled demographic seeds from an occupation and age-range table from the Labor Force Statistics[1].
For each occupation, the pipeline randomly selected an age range, sampled an age within that range, and assigned a gender from a fixed set of options.
These demographic seeds were used to prompt an LLM to generate structured personas containing a name, description, medications or supplements, general mood, and possible health conditions or injuries.
We then used… See the full description on the dataset page: https://huggingface.co/datasets/lbakar/health-log-extraction-dataset.cot-health-metrics-archive-2026-05-19
cot_health_metrics archive
Private archive of local /Users/ifc24/Develop/cot_health_metrics moved on 2026-05-19.
See MANIFEST.json and SHA256SUMS for source path, archive size, and checksum.
mental_health_counseling_responses
Dataset Card for Mental Health Counseling Responses
This dataset contains responses to questions from mental health counseling sessions.
The responses are rated by LLMs using the dimensions: empathy, appropriateness, and relevance.
A detailed explanation of the rating process can be found in this blog post.
For a detailed analysis of LLM-generated responses and their comparison to human responses, refer to this blog post.
The original data with the human responses can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_responses.gemma-4-E4B-it-Health_Benchmarks-benchmarkBenchmark of google/gemma-4-E4B-it against yesilhealth/Health_Benchmarks dataset.
Accuracy: 77.8%.
Metric
Value
Correct
5864
Incorrect
1669
Errors
2
Total samples
7535
Total completion tokens
8,144,545
Raw stats:
{
"accuracy": 0.778,
"correct": 5864,
"incorrect": 1669,
"error": 2,
"total": 7535,
"completion_tokens": 8144545
}
gpt-oss-20b-Health_Benchmarks-benchmarkBenchmark of openai/gpt-oss-20b against yesilhealth/Health_Benchmarks dataset.
Accuracy: 80.4%.
Metric
Value
Correct
6060
Incorrect
1475
Errors
0
Total samples
7535
Total completion tokens
2,940,179
Raw stats:
{
"accuracy": 0.804,
"correct": 6060,
"incorrect": 1475,
"error": 0,
"total": 7535,
"completion_tokens": 2940179
}
ehristoforu__Gemma2-9B-it-psy10k-mental_health-details
Dataset Card for Evaluation run of ehristoforu/Gemma2-9B-it-psy10k-mental_health
Dataset automatically created during the evaluation run of model ehristoforu/Gemma2-9B-it-psy10k-mental_health
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__Gemma2-9B-it-psy10k-mental_health-details.glp1-telehealth-rules-us
GLP-1 Telehealth Rules by U.S. State
Which U.S. jurisdictions (50 states + Washington, D.C.) require a live video
visit to start GLP-1 treatment by telehealth — plus each jurisdiction's medical
board, Medicaid GLP-1 coverage status, and nurse practitioner
prescriptive-authority classification (full/reduced/restricted) verified
against state nurse practice acts with statute citations.
Canonical, always-current source: https://www.pallashealth.co/glp-1/telehealth-rules
Live CSV:… See the full description on the dataset page: https://huggingface.co/datasets/pallas-health/glp1-telehealth-rules-us.Healthylife-insurance-charge-prediction-newhealthylife-insurance-charge-logMeera-Anilkumar-HealthyLife-Insurance-Charge-PredictionLLMgenerated_fictive_medical_report_and_summaries_with_omissions_label_Fr_Healthcare
🏥 French Synthetic Medical Reports and Summaries with Omission Labels (Fr-Healthcare)
This dataset contains fictitious French medical reports, each paired with a summary and a binary label indicating whether the summary omits relevant factual content. It is designed solely for evaluating factual consistency and omission detection in Natural Language Processing, particularly in the medical domain. We must emphasize that all names, identifiers, dates, medical information, and any… See the full description on the dataset page: https://huggingface.co/datasets/AchOk78/LLMgenerated_fictive_medical_report_and_summaries_with_omissions_label_Fr_Healthcare.zelk12__recoilme-gemma-2-psy10k-mental_healt-9B-v0.1-details
Dataset Card for Evaluation run of zelk12/recoilme-gemma-2-psy10k-mental_healt-9B-v0.1
Dataset automatically created during the evaluation run of model zelk12/recoilme-gemma-2-psy10k-mental_healt-9B-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__recoilme-gemma-2-psy10k-mental_healt-9B-v0.1-details.healthylife_insurance_prediction
